← All updates

Large Language Models Are Cultural Technologies

I have long been a fan of Alison Gopnik because of her research and writing about infants. She is also part of UC Berkeley's AI group. Along with Henry Farrell, Cosma Shalizi, and James Evans, she published what I consider to be an essential paper in Science. The article is paywalled, but Gopnik has published an image of the article on her site. This is compatible with my opinion that LLMs should be seen as tools for augmenting human intellect rather than autonomous intelligent agents (which they can emulate)

Debates about artificial intelligence (AI) tend to revolve around whether large models are intelligent, autonomous agents. Some AI researchers and commentators speculate that we are on the cusp of creating agents with artificial general intelligence (AGI), a prospect anticipated with both elation and anxiety. There have also been extensive conversations about cultural and social consequences of large models, orbiting around two foci: immediate effects of these systems as they are currently used, and hypothetical futures when these systems turn into AGI agents—perhaps even superintelligent AGI agents. But this discourse about large models as intelligent agents is fundamentally misconceived. Combining ideas from social and behavioral sciences with computer science can help us to understand AI systems more accurately. Large models should not be viewed primarily as intelligent agents but as a new kind of cultural and social technology, allowing humans to take advantage of information other humans have accumulated.

The new technology of large models combines important features of earlier technologies. Like pictures, writing, print, video, internet search, and other such technologies, large models allow people to access information that other people have created. Large models—currently language, vision, and multimodal—depend on the internet having made the products of these earlier technologies readily available in machine-readable form. But like economic markets, state bureaucracies, and other social technologies, these systems not only make information widely available, they allow it to be reorganized, transformed, and restructured in distinctive ways. Adopting Simon’s terminology, large models are a new variant of the “artificial systems of human society” that process information to enable large-scale coordination [(1), p. 33].

Our central point here is not just that these technological innovations, like all other innovations, will have cultural and social consequences. Rather we argue that large models are themselves best understood as a particular type of cultural and social technology. They are analogous to such past technologies as writing, print, markets, bureaucracies, and representative democracies. Then we can ask the separate question about what the effects of these systems will be. New technologies that are not themselves cultural or social, such as steam and electricity, can have cultural effects. Genuinely new cultural technologies—Wikipedia, for example—may have limited effects. However, many past cultural and social technologies also had profound, transformative effects on societies, for good and ill, and this is likely to be true for large models.

These effects are markedly different from the consequences of other important general technologies such as steam or electricity. They are also different from what we might expect from hypothetical AGI. Reflecting on past cultural and social technologies and their impact will help us to understand the perils and promise of AI models better than worrying about superintelligent agents.

That is, whether or not LLMs "think" or have some kind of understanding is not really relevant to their use as a tool. This article digs into that:

Human Intelligence, the Secret of Artificial Intelligence

In sum, the symbolic image, which is sensible and material, will trigger in the human mind the production and coherent weaving of an intelligible meaning from a multitude of semantic threads: a conceptual sense; a narrative sense through the reconstruction of syntactic trees and groups of paradigmatic substitutions; an intersubjective and social sense; an objective referential sense; an affective and memorial sense. That is to say that, once received by human intelligence, a material text becomes bound to an entire immaterial complexity, a complexity that is by no means random but rather strongly structured by languages, dialogue rituals and social rules, the logic of emotions, the contextual coherence inherent in corpora and worlds of reference. The capacity of language models to « reason » and respond to requests in a pertinent way is an effect of corpus, related to the priority given to dialogic training data and to data that adopt a demonstrative style. Enormous learning data enable a statistical capture of discourse norms.

Now it is precisely this solidarity between the material part of texts—now digitized—and their immaterial part that artificial intelligence will capture. Let us not forget that only the signifier (sequences of 0s and 1s) exists for machines. For them, there are neither concepts, nor narratives, nor subjects, nor worlds of real or fictional reference, nor emotions, nor resonances linked to personal memory, and even less any rooting in sensible experience of an animal type. It is only thanks to the gigantic quantity of training data and the enormous power of contemporary computing centers that statistical models manage to reify the relationship between the sensible form of texts and the multiple layers of meaning that a human reader spontaneously detects.

And this:

Rotating the Space: On LLMs as a Medium for Thought

Text as a Rotatable Space

Here’s a different way to think about it.

Language encodes ideas. But any particular text—a paragraph, an argument, an explanation—is not the idea itself. It’s a projection of the idea into a particular form. The same underlying concept can be expressed from different angles, at different levels of abstraction, for different audiences, through different metaphors, in different rhetorical modes.

Think of a three-dimensional object casting a shadow on a wall. The shadow is a two-dimensional projection. If you only see one shadow, you might mistake it for the thing itself. But if you can rotate the object—or equivalently, move the light source—you see different shadows. Each shadow reveals something about the object’s structure. No single shadow is the object, but multiple shadows from different angles let you reconstruct what the object actually is.

High-dimensional spaces work similarly, but with more complexity. A concept that exists in a thousand-dimensional space of meaning can be projected into the low-dimensional space of a particular text. That text captures some aspects and loses others. A different text—same concept, different projection—captures different aspects.

What LLMs enable is rapid rotation through projection-space.

You can take a half-formed idea and project it into the voice of a particular person. Then into a counter-argument. Then into a metaphor. Then into a dialogue between opposing positions. Then into an explanation for a child. Then into an explanation for a hostile expert. Each projection reveals different structure. Each rotation shows you something the previous view hid.

This is not summarization, which is more like dimensionality reduction—crushing the space down until it fits in a smaller container, losing most of the structure in the process. “Teach it to me like I’m five” is asking for PCA on a thousand-dimensional space, then wondering why you’ve lost nuance. Of course you have. That’s what flattening does.

What we’re describing is something different: using the LLM’s generative capacity to produce multiple projections, iteratively, so you can explore the structure of something too complex to see from any single angle.