# AI Doesn't Just Make Culture. It Interprets It.

2026-09-09 · Princeton, New Jersey · Reported Feature

A Princeton-linked workshop at one of machine learning's most important conferences is pushing a deceptively difficult idea: if generative AI operates inside language, art, context and meaning, then evaluating it cannot stop at whether the answer is technically correct.

A Princeton-linked workshop at one of machine learning's most important conferences is pushing a deceptively difficult idea: if generative AI operates inside language, art, context and meaning, then evaluating it cannot stop at whether the answer is technically correct.

---

For most of the modern artificial-intelligence boom, the public conversation has treated culture as something AI touches after the technical work is done. Engineers build the model. Companies release it. Then artists, teachers, journalists, policymakers and critics arrive to argue about bias, copyright, misinformation, labor, aesthetics and whatever the model has begun doing to the rest of us.

A group of researchers connected to Princeton University is proposing a different order of operations: culture is not the cleanup phase. It is part of the machinery.

On Sept. 2, Princeton's Center for Digital Humanities highlighted 'Culture × AI: Evaluating AI as a Cultural Technology,' a July workshop held at the International Conference on Machine Learning in Seoul, South Korea. Princeton described it as the first ICML workshop dedicated to culture and AI. The university said the event was organized by researchers from Princeton, the Alan Turing Institute, Google DeepMind and Cornell University, with Princeton English professor and digital-humanities director Meredith Martin among the organizers.

The distinction matters because ICML is not primarily a conference about cultural criticism. It is one of the major technical venues for machine-learning research. Putting a workshop about interpretation, aesthetics and cultural judgment inside that environment is an argument in itself: these questions belong upstream, while systems are being designed and evaluated, not only downstream after they have entered classrooms, studios, newsrooms and everyday conversation.

Culture is not another benchmark category

The workshop begins with a premise that sounds obvious until its consequences are taken seriously. Generative systems are trained on enormous quantities of human-made social and cultural material, then generate more cultural material: paragraphs, images, dialogue, summaries, scripts, code, video and increasingly combinations of all of them. Their output is not culturally neutral simply because it came from software.

The organizers argue that much of AI evaluation has approached culture defensively. Researchers test for offensive outputs, dangerous behavior, misinformation, bias or violations of stated values. Those are necessary questions. But they mostly define success as the absence of failure.

The Culture × AI workshop asks what a positive definition of cultural competence would look like. Could an AI system handle context well? Could it recognize that two interpretations may both be legitimate? Could it make an aesthetic judgment without reducing the judgment to a popularity score? Could it respond appropriately when meaning depends on history, social position, genre, irony, tradition or a community's internal vocabulary?

Those are not the kinds of problems that collapse neatly into a single correct answer. That is precisely the point.

The problem with asking a context machine for one right answer

A related research paper published in Frontiers in Artificial Intelligence in February gives the workshop's intellectual project a name: 'computational hermeneutics.' The paper, whose long list of authors includes Martin and several of the workshop organizers and speakers, describes generative AI systems as 'context machines' and argues that evaluating them requires more than standardized questions about accuracy.

Hermeneutics is the tradition of studying interpretation: how meaning is made, how context changes understanding and how readers or observers can arrive at different defensible readings of the same material. The researchers identify three recurring interpretive challenges for generative AI: situatedness, because meaning emerges in a particular context; plurality, because more than one interpretation can be valid; and ambiguity, because meanings can conflict without one side simply being an error.

That framework helps explain why ordinary AI conversations can become strange so quickly. Ask a model to calculate a percentage and there is usually a testable answer. Ask it what a joke means, whether a character is sympathetic, what a phrase implies in a particular relationship, why a song feels nostalgic or whether a piece of writing sounds sincere, and the task changes. The model is no longer retrieving a fact-shaped object. It is navigating meaning.

The authors are careful not to equate that process with human interpretation or to grant a model human intentions. Their point is more structural: the system must make context-sensitive selections to produce these outputs whether or not anyone wants to describe that process with human psychological language.

The papers make the abstraction concrete

The workshop's accepted presentations show how broad that territory already is. One paper examined whether narrative forecasting could measure tension in stories generated by language models. Another used LLM-based audience agents to simulate K-pop concert chat. Other presentations considered the cultural reach of generative systems, used narratology as a tool for evaluating model behavior, and proposed methodological principles for cross-cultural AI evaluation.

Taken together, those projects reveal the limitation of talking about AI only as a productivity tool. A system that helps draft an email may be a productivity tool. The same underlying class of system can also simulate a fan community, generate a story, characterize a person, explain a tradition, translate an idiom or tell a user what a piece of art supposedly means. At that point, it is participating in the circulation of culture whether the product label says so or not.

And participation is not the same thing as mastery. An AI can produce something that looks culturally fluent while flattening the thing it is describing. It can generate the recognizable surface of a genre while missing the reason people care about the genre. It can reproduce a community's language without belonging to that community, or confidently select one interpretation where a human reader would recognize unresolved ambiguity.

That creates an evaluation problem more difficult than catching a hallucinated date. A culturally bad answer can be factually composed of true sentences.

The humanities move upstream

The workshop's most consequential proposal may be organizational rather than philosophical. Its organizers want ideas from the humanities, arts and qualitative social sciences incorporated into the design and evaluation of AI systems earlier.

That reverses a familiar pattern in technology development. Humanists and cultural researchers are often invited to study what a system did after its core assumptions, datasets, interfaces and performance measures have already been chosen. Culture becomes an impact assessment.

The Culture × AI framing treats interpretation as a design problem. If a model is expected to operate in culturally complicated settings, then the ability to notice context, preserve uncertainty and accommodate legitimate disagreement cannot be an optional layer added after deployment.

Princeton is pursuing the same general question beyond the ICML workshop. Its Center for Digital Humanities is running 'Modeling Culture: New Humanities Practices in the Age of AI,' a year-long effort focused on how AI might contribute directly to humanities scholarship rather than only serving as an object of concern. The project acknowledges the familiar debates over bias, labor, environmental cost, intellectual property and education, but asks what happens if researchers also study the new methods AI makes possible.

That is a more complicated position than either AI boosterism or AI rejection. It assumes the technology is consequential enough to deserve criticism and useful enough to deserve serious methodological development.

Better cultural AI could also be more powerful cultural AI

There is an important tension embedded in the workshop's premise. Teaching machines to handle culture more effectively is not automatically a social good.

The organizers say so directly. Systems that become more sensitive to style, context, emotional cues, social norms and aesthetic preference could become more useful to artists and researchers. The same capabilities could make automated persuasion, synthetic entertainment, imitation and cultural substitution more effective. A model that understands the mechanics of belonging better might also become better at manufacturing the appearance of belonging.

The question therefore cannot be reduced to whether AI should become 'better at culture.' Better for whom, measured how, and toward what end are themselves cultural questions.

That is where the workshop's shift in vocabulary becomes useful. Technical evaluation traditionally rewards convergence: a model gets closer to the correct output. Cultural interpretation often requires the opposite instinct. Good judgment may mean recognizing that the evidence supports several readings, that the answer changes with context, or that uncertainty is part of the meaning rather than a defect to eliminate.

The machine is already in the conversation

For years, one of the central public questions about artificial intelligence has been whether machines can understand language. The Culture × AI project suggests a more immediate problem.

People are already using machines to help decide what language means.

They ask AI to interpret messages from partners and coworkers. They ask it to explain jokes, summarize books, judge tone, characterize historical events, recommend music, critique writing and generate images in recognizable cultural forms. Organizations use the same technologies to classify audiences, moderate speech and decide which pieces of culture receive attention.

Whether the machine 'understands' any of this in the human sense is still philosophically and technically contested. But the outputs enter human systems of meaning regardless. People read them, react to them, publish them, teach from them, make decisions from them and build more culture on top of them.

That makes evaluation harder than asking whether the model passed the test. In culture, the test is often part of what is being argued about.

Princeton's contribution to this conversation is not a declaration that AI has become a human interpreter. It is a warning that we may already be using interpretive technologies while evaluating them as if interpretation were merely another engineering feature.

If generative AI is going to help write, translate, recommend, explain and classify the materials through which people make sense of one another, then the humanities are not standing outside the technology looking in.

They are part of the specification.

SOURCE NOTES

• Princeton CDH - Sept. 2, 2026 announcement • Culture × AI - official workshop page • ICML 2026 - workshop announcement • Frontiers - Computational hermeneutics paper • Princeton CDH - Modeling Culture project

---

ProbleMattic is written and maintained by Matthew Kulcsar, a software engineer, project manager, technologist, platform builder, emergency-services-trained helper, grandfather, and lifelong collector of broken systems, odd behaviors, and useful nonsense.
