constructed startin with 1x1mm square; last square is 144x144
Published on

How to Ruin Credibility, Valuable Data, and Privacy Protections In One Summer

Authors
  • avatar
    Name
    Curtis Mitchell
    Twitter

Earlier this summer the US Department of Commerce issued directive DAO 216-26, an order which places severe restrictions on how the US Census Bureau and Bureau of Economic Analysis (BEA)1 add privacy protections to data that they routinely publish. This policy change impacts not only the privacy of individuals and businesses in the United States, but the trust they will have in government data collection going forward.

What DAO 216-26 Does

Most directly, the DAO directive states that "Noise infusion shall not be used for any statistical product." Noise infusion is a class of statistical methods that involves injecting random values into published statistics in a way that strikes a balance between data privacy and data usefulness (also known as "data utility"). The goal is to share data that is just accurate enough to be useful and also sufficiently inaccurate (or "noisy") to prevent people and organizations in the data from being reidentified by linking the same data values to external data sources2. In the place of noise infusion, the order requires that the Census Bureau and BEA use techniques known as "coursening" and "supression." Coursening is the rounding of data values by aggregating or grouping together values, while supression is outright removing or redacting certain values.

The Risks to Privacy

Noise infusion, especially the usage of differential privacy in the 2020 nationwide census, provides the strongest balance of privacy and data utility based on decades of developments in statistics and computer science. This order revokes the usage of differential privacy and even some weaker versions of noise infusion that were used in nationwide censuses from 1990 to 2010. Later research showed that some of those previously-used techniques, such as cell swapping (moving data values around so that it is unclear which data originally belonged to which entity) were still vulnerable to privacy violations3.

Returning to the usage of coursening and suppression moves Census Bureau and BEA statistics back to techniques that were standard in the 1970's, before the Internet, widespread data collection, and weekly data breaches made linking a person's census data to other datasets (and violating that individual's privacy in the process) much easier. My own work at the Census Bureau from 2023-2025 consisted of a number of pilot projects to test and demonstrate additional privacy technologies that the Bureau could implement to provide more privacy-protected data to researchers and collaborators. It seemed at the time that there was great enthusiasm for both providing high-quality data and robust privacy guarantees, but I fear that those goals are no longer a priority.

How this Harms Trust in Institutions

The disagreement over noise infusion is a highly technical debate that most US citizens and residents won't think about, at least not until the 2030 nationwide census. But DAO 216-26 has already harmed trust with academic researchers, business analysts, and others who use Census Bureau and BEA data on a regular basis. This order was published with no prior notice and without the usual period for public comment, preventing researchers and businesses from assessing and providing feedback on how it will impact their work.

Since its publication the Census Bureau itself has issued some guidance in a recently published Data Stewardship Program update (see section 7 specifically) and issued statements at the 2026 Joint Statistical Meeting in Boston earlier this month. But several questions remain unanswered, including which data publications will be affected going forward and how privacy risks will be formally measured under a return to using coarsening and suppression. Researchers are left wondering why this order was published without a clear explanation of its impact and how it will affect their work.

But Maybe Harms to Privacy Are the Real Goal

Last week a research paper was posted to the Census Bureau website that purports to analyze data about non-citizens voting in the the 2020 general election. The paper has been criticized for multiple irregularities such as having no listed authors and only a sparse description of its statistical methods. For example it has no explanation of how its authors linked census data to voter registration and dealt with sources of errors such as outdated naturalization records or duplicate names in voter registration data. When I published a working paper in 2024 during my time at the Census Bureau, my coauthors and I went through multiple rounds of reviews and edits for the paper despite the fact that it describes a proof-of-concept and did not include any published statistics. How this voting fraud paper passed the Bureau's own internal reviews and accuracy checks and why nobody signed their name to it is concerning for an institution that is meant to report truthful facts and figures about the country.

How does this new paper relate to the banning of noise infusion? Because implementing DAO 216-26 will reduce privacy protections in the 2030 nationwide census and thereby create a dataset of the US population that is more easily combined with other datasets, potentially ones with sensitive information such as voter registration and immigration status. Publishing statistics that appear legitimate and support a pre-existing political narrative would be easier than ever. And maybe that's the whole point of this policy order.

Further Reading & Watching

Footnotes

  1. While US citizens and residents are likely familiar with the Census Bureau, the Bureau of Economic Analysis is a similar organization that provides reporting on key economic indicators like the gross domestic product (GDP) as well as data on spending and income of the government, corporations, and individuals.

  2. Imagine a dataset that was used to describe a sensitive topic such reports of crime in a region. You'd want data that were accurate enough to be useful to understand different types of crimes being perpetrated, how effective policing is, etc., but not so detailed that someone could re-identify who specific victims of those crimes were.

  3. It should be acknowledged that some academic researchers and other users of these datasets were displeased with the usage of differential privacy in the 2020 census, often because it was felt that the prior privacy protections like cell-swapping were sufficient. However, recall that this order also restricts the usages of those other noise infusion techniques as well. Most researchers would acknowledge that coursening and supression are even more disfavored.

Enjoyed this? Get new posts by email or subscribe via RSS. No tracking, unsubscribe anytime.