In 2021, Roopika Risam discovered that the Social Security Administration didn’t know she was a U.S. citizen.
She immigrated to the U.S. as a child and became a naturalized citizen at 16, but no one had updated the record. For years she traveled internationally and worked without issue, unaware that a single wrong data point sat in a federal database with her name on it. Now, with immigration enforcement leaning harder on government data than ever before, it’s the kind of bureaucratic error that can upend a life.
That personal discovery anchors Risam’s new book, Data Empire: The Power of Information to Organize, Control, and Dominate. The professor of film and media studies and comparative literature traces data collection from Ice Age cave markings to the AI systems governments now depend on. From medieval land surveys to actuarial tables that valued enslaved people as cargo, the choice of what to count, and who counts, she argues, has always been about power.
The stakes aren’t abstract. Last year, Immigration and Customs Enforcement requested 1.3 million tax records from the IRS, a request a federal judge ultimately curtailed. As AI systems increasingly run on data infrastructure built and controlled by private companies, Risam’s history of who gets counted reads like a warning.
Risam’s research has drawn more than $4.3 million in grants from the National Endowment for the Humanities, the Mellon Foundation and others. She also created Torn Apart/Separados, a data visualization project that mapped the U.S. government’s family separation policy at the border, and is founding co-editor of Reviews in Digital Humanities and past president of the Association for Computers and the Humanities.
Ahead of its release, Data Empire made NPR’s and New Scientist’s July roundups of notable new releases.
In this Q&A, Risam discusses the ancient origins of data science, why the story of data has always been a story of power, and what she thinks citizens and policymakers must understand as AI reshapes who controls information today.
Why was it important to include your own experience with data, rather than keep the argument abstract?
Now, with immigration policy focused on finding, detaining, and deporting people on the basis of their categorization in government databases, the error feels chilling. One wrong data point could upend your entire life. And we’ve seen this happening to people; it’s rare but still, citizens have been mistakenly deported.
As Immigration and Customs Enforcement partners with other government agencies to access their data, hundreds of millions of people are vulnerable. Just last year, ICE partnered with the IRS to access their records. By law, they should only be able to access IRS data for federal criminal investigations. But ICE requested 1.3 million records. Fortunately, a federal judge issued an injunction, ruling that the data should only be shared with a government official involved in a federal case and that an ICE agent requesting tens of thousands of people’s information obviously wasn’t investigating all those people personally.
My experience at the Social Security Office made it clear that anyone can be at risk without even knowing it and, therefore, we all need to take the power of data and its potential for material effects on our lives seriously. Other people are realizing this, too. I think, for example, of the people of Delaware who successfully demanded that the state’s motor vehicle registries stop providing data to ICE. Once we realize what’s happening, we become empowered to do something about it.
You argue that humans are “the data species” and trace data collection back tens of thousands of years. What surprised you most about the origins of data, and how did researching those early forms of recordkeeping change the way you think about today’s digital world?
When I began working on the book, I intended to write about the relationship between data today and European overseas empires of the past. Even with that scope, “data” was older than most data histories suggest. But as I was doing research, it was obvious that it didn’t suddenly appear out of nowhere in the 15th century. I needed to know just how far back it went, so I kept looking for examples in archives and museums and it took me further back into the medieval period, then the beginning of the Common Era, and I just kept going. I assumed I’d be able to stop where writing emerged, around 3400 BCE or so, but I was wrong about that. Writing itself was only possible because of data. The cuneiform script developed in Mesopotamia from simple marks on clay spheres used to store tokens: simple clay shapes used to represent goods. These tokens were essential to the rise of civilization itself, as they made managing the scale of resources needed to feed increasingly more dense populations possible. But it didn’t stop there!
The cave art in Lascaux, France, from approximately 20,000 years ago includes dashes, dots, and y-shaped marks hypothesized to track animal breeding cycles so the people living there could identify the best times to hunt for long-term sustainability of game.
And as early as 43,000 years ago, ancient peoples were carving notches into baboon bones to keep track—of what, of course we don’t know. But it’s incredible to see how integral creating data has been to humans for so long; it’s one of the distinguishing features of our species.
As a result, the way I think about the digital world completely changed. We often talk about data as though it’s something that only exists because of scientific practices developed in Europe in the 18th century or the rise of computation in the 20th century. But humans have been creating data for millennia. We have to ask ourselves what’s actually new now? It’s the scale, speed, and concentration of power that machines make possible. Seeing that long history helped me realize that today’s debates about AI, surveillance, and the tech industry are really the latest chapter in a much older story about how humans use information to organize society and, of course, who gets to control it.
A central theme of Data Empire is that data has always been intertwined with power. When did you realize that the story of data was also a story about governance, control, and inequality, rather than simply technological progress?
I noticed a pattern that kept appearing, regardless of the time or place I was studying. When I looked at the case of ancient Mesopotamia, I saw that when token storage began to be concentrated within temples, the administrative centers of the earliest civilizations, priests assumed a central role in political power, some becoming priest-kings. In the Levant, around the beginning of the Common Era, I found a community of Jews, led by Judas of Galilee, refusing to let the Roman Empire count the people of Judaea by census; they recognized it as a violation of their sovereignty and direct contradiction of instructions in Exodus. Within 20 years of the Norman Conquest in 1066, one of the first things William the Conqueror did was order a massive land survey to record who owned plots of land and legitimize Norman presence there. And I kept on going—European empires using censuses, land surveys, and studies of people to justify their rights over lands in the Americas and the Indigenous peoples living there, data about human bodies justifying “scientific” theories of white supremacy.
We also saw the emergence of national quantification in the 20th century as the World Bank and IMF decided whether to invest in new nations that fought off colonialism, with the goal of remaking their economies in the mold of liberal capitalism. The trend is undeniable. Data has long raised the very questions we must ask of it today: Who gets counted? Who decides which kinds of data matter? Who benefits from collecting it? And, of course, who bears the consequences?
Many people think of data as objective or neutral. In the book, you challenge that assumption. What are some of the most important ways that human choices shape data and the decisions made from it?
One of the biggest misconceptions about data is that it simply reflects reality. But data has always been the product of human choices. We decide what to measure, what categories we use to describe it, what counts as evidence, which data survives, which disappears. Those decisions shape the stories data can tell long before a person or, now, an algorithm even analyzes it.
History is full of examples. Colonial governments created censuses to sort people into racial and religious categories to govern them more effectively. Insurance companies like Aetna developed actuarial tables to value human lives—and even insured enslaved Africans as “cargo” during the Transatlantic Slave Trade and on plantations in the U.S. South, compensating policy holders in case of death or self-emancipation. Libraries in the U.S. and Europe built organizational systems that privileged some forms of knowledge and marginalized others, starting with the Dewey Decimal System and then the Library of Congress Subject Headings. Today’s AI systems inherit many of those same dynamics because they are trained on data created by people and institutions with their own priorities and biases.
This doesn’t mean that data is inherently evil or useless or we shouldn’t trust it. Data is one of humanity’s greatest inventions. It has allowed us to build cities, understand disease, explore space, make extraordinary scientific discoveries, and, most importantly, tell our own stories. But data is never self-explanatory. We need to always be asking who created it and for what purposes, whose experiences it captures and whose it overlooks. Those are just as much historical and ethical questions as they are technical ones.
Dartmouth has a unique place in the history of computing and AI through figures such as John Kemeny and Thomas Kurtz. How does Dartmouth’s role in the development of modern computing fit into the larger story you tell in Data Empire?
I couldn’t miss the opportunity for a Dartmouth shoutout in Chapter 9! I include the story of the Homebrew Computer Club, a group of engineers, hobbyists, and countercultural tinkerers who gathered in a Menlo Park, California, garage in the 1970s to share code and imagine a more democratic future for computing. One of the technologies that helped make personal computing possible was BASIC, the programming language developed at Dartmouth by Kemeny and Kurtz. BASIC was designed to make programming accessible to Dartmouth students and ultimately, everyone, rather than a small group of specialists. Their commitment to widening participation in computing fit perfectly with the ethos of Homebrew.
This is also where I got to throw a little shade at Bill Gates. Gates and Paul Allen built Altair BASIC on top of the BASIC language, which Kemeny and Kurtz had made openly available. But when copies of Altair BASIC began circulating within Homebrew, Gates famously responded with his “Open Letter to Hobbyists,” criticizing people for sharing software without paying for it. The irony! It’s a fascinating moment because it captures a larger transition I trace in Data Empire: the shift from computing as a shared commons to computing as a commodity.
Of course, this year, we’re celebrating the 70th anniversary of the 1956 Dartmouth Summer Research Project on Artificial Intelligence, where the term “artificial intelligence” was coined. Now, as I suggest in Chapter 10, AI companies are gaining extraordinary influence over core functions of modern states as governments increasingly rely on them for data management services.
Dartmouth sits at the center of two important moments in computing history, and together they illuminate one of the central themes of Data Empire: technologies are not inevitable. They are the product of human choices about who gets access, who exercises control, and whose interests innovations should serve.
The book concludes by looking toward the next century. As AI becomes more deeply embedded in everyday life, what lessons from the history of data do policymakers, technologists, and ordinary citizens today need to understand?
Right now, we’re being told that AI is inevitable, that we need to get on board or get left behind. It’s as if this technology and the Big Data that fuels it simply materialized or are a logical outcome of progress. But one of the central lessons of Data Empire is that this isn’t true. AI is the product of a long series of human choices: to collect ever more data, to allow companies to accumulate unprecedented amounts of it, and, increasingly, to rely on those companies to provide cloud infrastructure, data storage, surveillance systems, and automated decision-making for governments.
This should give us pause. Throughout history, states have often partnered with private companies to extend their reach. A consequential example is the British East India Company, which began as a commercial enterprise before taking on many of the functions of government on behalf of the British Empire. Another is the partnership between IBM and Nazi Germany, which made identification of Jews more efficient with tabulating machines and enabled the Holocaust. Today, we continue to witness private companies becoming deeply embedded in the infrastructure through which governments operate.
As governments forge public-private partnerships with tech companies such as Google, Palantir, and Amazon, they are handing immense amounts of power to tech companies. This is the biggest threat to democracy right now, that private actors largely immune to government oversight, have their hands all over the data infrastructures that run countries. The lesson of history isn’t that we should reject new technologies. It’s that we shouldn’t treat their current trajectory as inevitable. It was the product of a series of choices and our leaders can choose differently—and we must hold them accountable for that.
People are beginning to realize this and some governments are, too. For example, the Dutch Ministry of Defense recently announced that they don’t want to partner with Palantir anymore; they don’t want to be dependent on the company. Other countries in Europe are leaning that way, too. Now is a crucial moment to rethink who controls our data, how these technologies are regulated, and whether they strengthen democratic institutions or concentrate power in the hands of a few corporations.
Join Roopika Risam at a Dartmouth Libraries Book Talk in Baker-Berry Library on Tuesday, Sept. 29, at 12:30 p.m. to hear more about Data Empire. (Registration is required.)