How American AI Companies Are Colonizing Global Cultural Heritage
Silicon Valley's “information wants to be free” ethos treats cultural knowledge as a natural resource awaiting capture and optimization. Dissenting views are dismissed as backwards protectionism.
We commonly hear that when OpenAI, Anthropic, and Google trained their frontier AI models, they “scraped the Internet.” But this euphemistic framing strategically obscures the more pernicious reality that what these companies actually did was furtively extract cultural knowledge from across the globe, including Korean traditional painting techniques, Arabic calligraphy styles, Indigenous ceremonial imagery, African textile patterns, and much more. All of that human art and knowledge, now stored in Silicon Valley, was “lifted” (stolen) without permission, compensation, or acknowledgment in order to build privatized, large language model (LLM) systems for profit. Today, the databases powering American AI systems contain billions of images, texts, and audio recordings that represent centuries of non-Western cultural production, and while it is easy to point to the issue of copyright infringement here, what we are arguably faced with is a new form of cultural imperialism in which American technological dominance enables the mass appropriation of cultural heritage under the guise of “training data.”
Consider the LAION-5B dataset, which trained Stable Diffusion, one of the most widely used generative AI models. Researchers analyzing its contents found a massive overrepresentation of Western imagery alongside the systematic extraction of non-Western cultural materials which include, among other things, museum collections from Seoul and Cairo digitized for public access, now training commercial AI; sacred Indigenous imagery from communities that never consented to digital reproduction; and contemporary artists from the Global South whose Instagram posts became training data for systems that could replicate (and devalue) their labor.1 Unsurprisingly, the economics of this situation are scandalously lopsided because American AI companies like OpenAI (valued at over $80 billion) and Anthropic (valued at $18 billion) basically capture all the value while international communities whose knowledge trained these systems receive nothing whatsoever. When Korean artists organized protests in Seoul last year after discovering their work in various AI training datasets, OpenAI’s response was as ridiculous as it was telling—the company claimed that training on publicly available data is “transformative use” protected under American law. But why should American legal frameworks govern the appropriation of Korean cultural production? This absurd premise was simply taken for granted.2
There is a comparison to be observed here with British courts that validated the appropriation of colonized lands and resources through doctrines like terra nullius. At present, American tech companies like OpenAI are content to invoke their alleged right to “fair use” and “data scraping” in order to legitimize cultural appropriation at a mass scale. It is hard to see how this is fundamentally different from the structures of colonial metropolitan law, which expediently justified resource extraction from the peripheries according to its own dictates (and, certainly, with no concern for those directly affected or harmed). Perhaps the only difference is speed and scope; extraction and theft by gunboats and a complex administrative machinery is a lot harder than by server farms and web crawlers.3
What makes this extractivism distinctly American is not just that most AI companies are U.S.-based (they are), but how deeply it reflects American ideologies about information, property, and cultural production. Silicon Valley’s “information wants to be free” ethos, which stems from California counterculture and libertarian economics, treats cultural knowledge like a natural resource waiting to be captured and optimized. That different societies might have alternative frameworks for governing cultural materials is simply dismissed as backwards protectionism.4 So the troubling implications extend beyond the discrete economics of AI extraction; there is also a moral dimension to be considered in the question of cultural sovereignty, or the right of communities to control their own knowledge and determine how it circulates. When Stability AI trained models on images from the Smithsonian’s collections of Native American ceremonial objects, for example, they violated not just copyright but Indigenous protocols governing sacred knowledge transmission. Items usually restricted from photography in physical museums are now freely reproducible through AI, and there is no mechanism for communities to enforce traditional governance as a preventative measure.5
Through my own work at South Korea’s National Assembly, and in organizing international AI-governance forums, I have similarly observed how Korean artists working with traditional minhwa (folk painting) are despairing that they can do nothing about the fact that their techniques (which are passed down through generations and refined over decades of practice) are being reproduced by AI systems that anyone is able to access for $20 a month. Across Asia, artists and creators are realizing they are in a similar bind, including Chinese calligraphers and Japanese ukiyo-e artists who are finding that their distinctive styles are being replicated instantly through basic text prompts. There is some resistance emerging, but it is incomplete. South Korea’s 2024 Act on the Promotion of the Arts, which I helped research during my time at the National Assembly, establishes legal infrastructure for collective cultural data governance. The law mandates authentication systems, creates resale rights requiring royalty payments, and establishes a national art database, thus establishing the foundations for what we call “cultural data trusts.”6
Data trusts enable communities to pool their data rights and exercise them collectively by creating bargaining power against tech giants. This simply means that instead of individual artists negotiating separately with OpenAI or Anthropic—David versus Goliath at a global scale—cultural data trusts negotiate on behalf of entire communities in order to secure fair compensation, attribution requirements, prohibited uses, and ongoing monitoring. They represent a fundamentally different model from American platforms’ “terms of service” approach where users either accept extraction or are ignored and excluded.7
Outside of South Korea and Asia, Indigenous communities worldwide are also helping to lead this resistance. The Local Contexts initiative developed by Māori and First Nations communities, fo example, creates Traditional Knowledge Labels and Biocultural Labels that assert community governance over digital cultural materials. These labels explicitly challenge Western intellectual property frameworks embedded in American AI development by insisting that cultural knowledge belongs to communities and not corporations or “the commons.”8 Canada’s approach represents yet another path that emphasizes rebalancing power and recognizes that when individual creators negotiate with American platform monopolies, structural inequality makes genuine consent impossible. Their proposed framework would require AI companies operating in Canada to license training data through registered cultural data trusts or face significant penalties. These are just some ways impacted communities are trying to assert national sovereignty over their cultural governance against American tech dominance.
All of this is to say that the struggle over AI training data should be seen in the broader context of American power in the twenty-first century if it is to be properly understood. If military and financial institutions largely underpinned and defined U.S. dominance in the twentieth century, there is a case to be made that, today, that power increasingly operates through data extraction and infrastructure. Consider the fact that the Pentagon’s global base network is mirrored by Amazon Web Services’ data centers, or that the International Monetary Fund’s structural adjustment programs find a parallel in “terms of service” that frame and define creative labor worldwide. All of this creates serious challenges for the prospect of resistance, for how exactly do you pursue the nebulous undertaking of “resisting” structures you depend on daily? For example, Korean artists need Instagram to reach global audiences, but Instagram’s parent company, Meta, uses their posts to train AI. And many museums the world over must digitize collections for preservation and access, but as the above examples show, digitization clearly makes them vulnerable to extractive scraping.
Some recent, preliminary governmental responses to this problem include the EU AI Act, which demonstrates how regulatory power can constrain American tech companies by requiring transparency about training data and establishing accountability for AI harms. Similarly, China’s approach (though problematic in other dimensions) shows how states can build alternative AI ecosystems not dominated by Silicon Valley. Additionally, grassroots movements worldwide are creating peer governance structures that bypass American platforms entirely.9 These are all positive efforts, but the question remains as to whether these forms of resistance can coalesce into systematic alternatives or if they are destined to remain fragmented and therefore easily absorbed by platform capitalism. American AI companies have proven adept at co-opting critique (think “ethical AI” initiatives, “responsible innovation” rhetoric, and diversity programs) while maintaining extractive structures, and breaking this pattern will require more than “better regulation,” as it were, but entirely different imaginaries about how technology, culture, and power should relate.
From Seoul, I restlessly watch how American AI development is reshaping global cultural production. The generative models trained on scraped data are morphing things rapidly. When AI can generate passable Korean traditional painting, Indonesian batik patterns, and Persian miniatures instantly and cheaply, what happens to artists who spent decades mastering these traditions? How do non-Western artistic traditions survive? These are questions of civilizational import that cannot be ignored because if American companies control the infrastructure determining how cultural knowledge is stored, accessed, and reproduced, they effectively control the conditions of cultural possibility globally. Resisting this requires asserting cultural sovereignty in digital contexts. This includes the right of communities to govern their own cultural knowledge, to determine acceptable and prohibited uses, and to participate in technological development as partners rather than data sources. The cultural data trusts emerging in South Korea, Canada, and Indigenous communities worldwide represent one model for this sovereignty, but they remain imperfect, experimental, and face ongoing challenges. Still, they demonstrate that American AI companies’ extractive model is not inevitable, and that with institutional innovation and political will, alternative futures remain possible.
The story of America in the twenty-first century is inseparable from the story of American technology companies reshaping global culture. Tackling the new American extractivism head-on has to happen if we are going to imagine and create more equitable digital futures. What we should be chiefly and collectively thinking about is less whether AI will transform cultural production (it will), and more whether that transformation will serve diverse global communities or continue to concentrate power and profit in Silicon Valley. Cultural data trusts offer one path toward the former. Whether we take it is a political choice we have to make now.
Analysis of LAION-5B dataset: Abeba Birhane et al., “Multimodal Datasets: Misogyny, Pornography, and Malignant Stereotypes,” arXiv preprint arXiv:2110.01963 (2021).
OpenAI’s legal position on training data: Pamela Samuelson, “Generative AI Meets Copyright,” Science 381, no. 6654 (2023): 158-161.
On legal frameworks legitimizing colonial extraction, see Antony Anghie, Imperialism, Sovereignty and the Making of International Law(Cambridge University Press, 2005).
On Silicon Valley ideologies and information governance, see Fred Turner, From Counterculture to Cyberculture: Stewart Brand, the Whole Earth Network, and the Rise of Digital Utopianism (University of Chicago Press, 2006).
Indigenous perspectives on AI and cultural sovereignty: “Incorporating Indigenous Knowledge Systems into AI Governance: Enhancing Ethical Frameworks with Maori and Navajo Perspectives,” Preprints.org (December 2024).
Mika Noh, “Globalizing Art Law: South Korea’s Art Promotion Act as a Model,” Harvard Art Law Organization (August 2025); Mika Noh, “Art Promotion Act: Concerns and Expectations,” National Assembly Research Service (June 2024).
On data trusts: Sylvie Delacroix and Neil D. Lawrence, “Bottom-up Data Trusts: Disturbing the ‘One Size Fits All’ Approach to Data Governance,” International Data Privacy Law 9, no. 4 (2019): 236-252.
Jane Anderson and Kimberly Christen, “Decolonizing Attribution: Traditions of Exclusion,” Journal of Radical Librarianship 5 (2019): 113-152.
European Commission, “Artificial Intelligence Act,” Official Journal of the European Union (June 2024).




