AI for IA: large language models as thinking tools

tl;dr Asking "are LLMs sentient?" is not as valuable as "how can LLMs extend human cognitive and communicative capabilities, and how can we avoid this new cultural technology's pitfalls?
Tools to Think With
I've been thinking about the use of technology to augment human intellect for most of my life. Large Language Models appeal to me more as amplifiers of human intellect and knowledge creation than as artificial intelligences.
My 1968 Reed thesis (PDF) was about the use of neurofeedback as both an objective probe of the previously unquantifiable phenomenology of consciousness and as a first step toward a technology of consciousness manipulation more precise than psychedelics.
I chose the title "Tools for Thought" to my 1985 "History and Future of Mind Amplifiers" one afternoon in the library of Xerox PARC in the early 1980s when I came across a sentence in a publication about graphic user interfaces to the effect that "the screen serves as a visual cache for memory," thus extending the capacity of immediate memory beyond the "Seven plus or minus two: Some limits on our capacity for processing information."
In writing about the origins of personal computing, I came across Doug Engelbart's paper that has inspired so many, "Augmenting Human Intellect." I sought out Engelbart, interviewed and wrote about him. I taught his paper and he guested in my Stanford class. We became friends. I even made a brief video of the time Doug and his wife and Ted Nelson and his wife came to dinner at our place.
I continue to think that what Doug had in mind was something far beyond the mouse and graphical interface: his insistence on a framework of interoperating "Humans using Language, Artifacts, Methodology, and Training" seems particularly appropriate when applying Large Language Models: the way the querent uses language shapes the way LLMs respond; methodologies of human-ai discourse are evolving -- an encounter with a chatbot today is more like a conversation than traditional search querying -- and it is already clear that while the cost of access to LLMs is likely to drop, a gap will open between those who learn (are trained) how to use them to their benefit -- and avoid the pitfalls and hallucinations that await uncritical use of a tool that can be fatally wrong -- and those who do not know how to use this new cultural tool. Literacy lags technology. At certain levels of complexity, training has to be part of the technology.
In recent years, Cassandra Xia and Patricia Maes expanded on the artifact part of the framework:
Fifty years ago, Doug Engelbart created a conceptual framework for augmenting human intellect in the context of problem-solving. We expand upon Engelbart's framework and use his concepts of process hierarchies and artifact augmentation for the design of personal intelligence augmentation (IA) systems within the domains of memory, motivation, decision making, and mood. This paper proposes a systematic design methodology for personal IA devices, organizes existing IA research within a logical framework, and uncovers underexplored areas of IA that could benefit from the invention of new artifacts.
In 2019, Andy Matuschak and Michael Nielsen wrote at length about "How can we develop transformative tools for thought?" and they noted what Engelbart foresaw and manifested: Designers of thinking tools can use those tools to build better and more powerful thinking tools. Engelbart called it "bootstrapping."
The musician and comedian Martin Mull has observed that “writing about music is like dancing about architecture”. In a similar way, there’s an inherent inadequacy in writing about tools for thought. To the extent that such a tool succeeds, it expands your thinking beyond what can be achieved using existing tools, including writing.
The more transformative the tool, the larger the gap that is opened. Conversely, the larger the gap, the more difficult the new tool is to evoke in writing. But what writing can do, and the reason we wrote this essay, is act as a bootstrap. It’s a way of identifying points of leverage that may help develop new tools for thought.
So let’s get on with it. "
Written before Large Language Models set off the LLM co-evolution, Matuschak and Nielsen dove deeply into principles of thinking tool design, specifically concentrating on the mnemonic aspect -- expanding human memory in the tradition of Vannevar Bush's Memex.
Bret Victor's 2012 speech "inventing on principle" inspired many designers and specifically touched on an aspect of thinking tools that is key to the power of LLMs: tools that provide immediate feedback during the creative process can profoundly enhance creativity and problem-solving. (See "bootstrapping.")
In other words, AI can be used to empower IA, and some designers are already thinking and building, but AI for IA won't happen on a large scale without conscious, concerted effort -- first we need to learn how to design and spread informed use of the human, artifactual, linguistic, learning aspects of LLMs as psycho-techno-social systems. Just as spreadsheets and video games were early use cases for personal computers, YouTube, Wikipedia, and the Web emerged only when millions of people started to explore what the new media were capable of doing. Writing term papers and emails are contemporary commercial use cases for LLMs, but they are not the capabilities that are likely to elevate the noosphere.
The Extended Mind
I've kept my antennae tuned for signals of progress in AI for iA. In 2011, I heard that Andy Clark, a professor of cognitive philosophy -- which already sounds like the right track for what I seek -- was talking about the "extended mind." When challenged to demonstrate his assertion that much mental work involves external tools, he asked the challenger to multiply two four-figure numbers in his head. I interviewed Clark in 2011. Clark, together with David Chalmers, first published about the extended mind in 1998.
My interest in extended mind theories as a guide for design of mind amplifiers was invigorated by Annie Murphy Paul's book The Extended Mind: The Power of Thinking Outside the Brain:
Many tomes have been written on human cognition, many theories proposed and studies conducted (Tversky’s and Kahneman’s among them). These efforts have produced countless illuminating insights, but they are limited by their assumption that thinking happens only inside the brain. Much less attention has been paid to the ways in which people use the world to think: the gestures of the hands, the space of a sketchbook, the act of listening to someone tell a story or the task of teaching someone else. These “extra-neural” inputs change the way we think; it could even be said that they constitute a part of the thinking process itself.
Inspired by that book, I've written twice before about extended mind theory and practice: First post; Second post. (As a former writer of narrative non-fiction, I applaud the way Murphy Paul provided an entertaining literature review of extended mind research, in a sense creating an interdisciplinary field of inquiry that had not existed before she demonstrated the convergences of research in fields as diverse as gestures, workplace design, collaborative milieu ("scenius".)
Effective and humane design of mind amplifying tools requires a broader interdisciplinary approach than the engineering-heavy side of the design discourse. I have learned that any kind of interdisciplinary initiative takes time. In 2005, my TED talk called for an interdisciplinary study of cooperation. It took ten years for ASU Interdisciplinary Cooperation Initiative to be founded.
In 2012, I published an ebook, Mind Amplifier, specifically as an interdisciplinary curriculum for the designers of cognitive tools. Engelbart is essential, but so is the co-evolution of human culture and communication media and the work of Elinor Ostrom (institutions for collective action and governing the commons), Ivan Illich (convivial tools), Walter Ong, (orality and literacy), Lewis Mumford (the myth of the machine), Elizabeth Eisenstein (the printing press as an agent of change), Denise Schmandt-Bessaret (how writing came about), Robert K. Logan (the alphabet effect), and others outside the worlds of computation, telecommunication, and software engineering.
My own intellectual and tool-using trajectory was heavily influenced by Alan Kay's 1977 "Microelectronics and the Personal Computer," which led me to PARC, where Kay had been one of the most important drivers of development of the graphical user interface. I was thrilled by the rumor that it was possible to type and edit words on screens (instead of retyping pages with typewriters), and Kay's article showed that it was a reality in one R&D hive on Coyote Hill Road.
Kay, who certainly had the idea of computers as thought-tools in mind, helped turn the computer from an instrument controlled by specialists through the invocation of artificial programming languages into a thinking and communicating tool that the entire population can use by pointing and clicking on graphical depictions of familiar objects: the desktop metaphor enabled a user to find or store or open or discard a file by manipulating the graphic symbols for files, folders, trashcans. The GUI provided broad access to the personal computer -- previously used solely for scientific calculations, business data processing, and video games -- as a tool for thinking.
Cultural Technology: Better Frame than Artificial Intelligence?
One theme that has emerged for me in the developing narratives about artificial intelligence is that large language models and their chatbots can most productively be thought of as thinking tools -- cultural technology -- partners with rather than artificial replacements of human intellect. One foundational document of the development of the personal computer and digital networks was J.C.R. Licklider's 1960 Man-Computer Symbiosis:
Man-computer symbiosis is an expected development in cooperative interaction between men and electronic computers. It will involve very close coupling between the human and the electronic members of the partnership. The main aims are 1) to let computers facilitate formulative thinking as they now facilitate the solution of formulated problems, and 2) to enable men and computers to cooperate in making decisions and controlling complex situations without inflexible dependence on predetermined programs. In the anticipated symbiotic partnership, men will set the goals, formulate the hypotheses, determine the criteria, and perform the evaluations. Computing machines will do the routinizable work that must be done to prepare the way for insights and decisions in technical and scientific thinking.
(NB: "set the goals, formulate the hypotheses, determine the criteria, and perform the evaluations" also describes a good prompt for an LLM.)
Licklider was not just prescient: he was in large part responsible for the US Defense Department's Advanced Research Projects Agency's support for the development of personal computers and computer networks. The use of computer networks for social communication was also foreseen and facilitated by Licklider and Taylor in 1968. Time and again, the real magic powers of digital tools were not unlocked by the most prominent scientific, military, or business use cases. It took Licklider, Engelbart, and Taylor and their teams at ARPA and PARC to create what their legendary engineers wanted for themselves: mind amplifiers capable of designing better mind amplifiers.
Engelbart, Licklider, Taylor, Kay manifested today's technologies by looking beyond the immediately practical use cases through the eyes of visionaries (Engelbart), scientists (Licklider) and programmers who were also educators and musicians (Kay). Currently Alison Gopnik, an expert in early childhood learning, articulates the argument that LLMs are best seen as cultural tools that extend human cognitive capabilities rather than agents that imitate or replace them:
A common way of thinking about these systems is to treat them as individual agents, and then debate how intelligent those agents are. The phrase “an AI” rather than “AI” or “AI system,” implying individual agency, is frequently used. Some have claimed that these models can tackle complex commands, perform abstract reasoning, such as inferring theory of mind.
We argue that this framing is wrong. Instead, we argue that the best way to think of these systems is as powerful new cultural technologies, analogous to earlier technologies like writing, print, libraries, internet search and even language itself. Large language and vision models provide a new method for easy and effective access to the vast amount of text that others have written and images that others have shaped. These AI systems offer a new means for cultural production and evolution, allowing information to be passed efficiently from one group of people to another .
I have written about Henrich et al's conception of the importance of the transmission of ideas via cultural evolution -- the way useful ideas spread through populations and over time. Thinking occurs not just in part outside the brain, it also occurs largely outside the individual. Culture is the accumulation of what humans as a species learn and teach, which scaffolds what individuals know and do. And culture doesn't stop at accumulation -- it evolves, building on prior capabilities, forcing the brain to learn to read by exapting fight-or-flight mechanisms such as fast understanding of simple images (like tracks in mud ) and neural sequencing (that evolved for throwing stones at moving targets) into a new capability of visually transmitting knowledge across time and space.)
I asked chatGPT to summarize Alison Gopnick's conception of LLMs as "cultural technologies" -- tools, rather than agents:
Tools, Not Minds:
Gopnik argues that LLMs are not akin to human minds; they lack the intrinsic cognitive processes, intentions, or consciousness of a human agent.
Instead, they are advanced statistical tools trained on vast datasets to generate responses based on patterns in human language.
Cultural Artifacts:
LLMs are products of collective human knowledge and culture, much like books, maps, or other tools that externalize and share information.
They reflect and amplify existing cultural practices but do not independently innovate or possess intrinsic creativity.
Amplifiers of Human Intelligence:
LLMs act as extensions of human cognition by enabling people to process, synthesize, and communicate information more efficiently.
They serve as instruments for collaboration, exploration, and problem-solving, augmenting human potential rather than replacing it.
Importance of Human Oversight:
As tools, LLMs depend on humans for guidance, interpretation, and ethical use.
Gopnik highlights the risks of anthropomorphizing LLMs, cautioning against treating them as agents with independent judgment or responsibility.
Cultural Evolution:
LLMs represent a new phase in cultural technology, building on the tradition of tools like writing or printing that have transformed how humans think, learn, and communicate.
They enable access to and reconfiguration of knowledge in ways that can enhance collective intelligence and cultural development.
By framing LLMs as tools rather than agents, Gopnik stresses the importance of using them responsibly to complement and extend human capabilities, while remaining mindful of their limitations and the cultural values embedded in their design and deployment.
A number of other thinkers are beginning to communicate about LLMs as thinking tools, or -- as chatGPT told me -- "external cognitive scaffolds." The following excerpts are part of a longer and richer piece by John Nosta on "Large language models and the path to the higher self." (See also Nosta's "Are LLMs the New Cognitive Optimizer?")
The idea that tools augment our humanity is not new. Eyeglasses enhance sight; language extends thought; art expresses what words cannot. Far from dehumanizing us, such technologies deepen our engagement with the world and with ourselves. They are, in a sense, extensions of our highest aspirations—truth, beauty, and goodness—ideals central to Aristotle’s philosophy of human achievement.
...
What emerges in conversation with an LLM is what we can call a "cognitive dance"—a dynamic interplay between human and artificial intelligence that creates patterns of thought neither party might achieve alone.
Apropos of Gopnik's reference to cultural evolution, In 2018, Cecilia Hayes wrote about Cognitive gadgets: The cultural evolution of thinking. Culture is key to the use of LLMs as cognitive tools because culture is not designed or programmed (although design and programming may be its subject matter) but learned, communicated, expanded, and evolved through practice, teaching, and learning. Again, these essential parts of human augmentation tools are neither hardware, nor software, but learned and evolved cultural practices.
Précis of Cognitive Gadgets: The Cultural Evolution of Thinking
Cognitive gadgets are distinctively human cognitive mechanisms – such as imitation, reading, and language – that have been shaped by cultural rather than genetic evolution. New gadgets emerge, not by genetic mutation, but by innovations in cognitive development; they are specialised cognitive mechanisms built by general cognitive mechanisms using information from the sociocultural environment. Innovations are passed on to subsequent generations, not by DNA replication, but through social learning: People with new cognitive mechanisms pass them on to others through social interaction. Some of the new mechanisms, like literacy, have spread through human populations, while others have died out, because the holders had more students, not just more babies.
A Tool That Teaches How To Use It
In a sense, the GUI is a thinking tool that teaches people how to use it. To be sure, there are levels of expertise that knowledgeable engineers (and teenagers) apply beyond the prosaic world of files, folders, and trash cans. The nested abstractions of operating system and machine language are hidden from the everyday user by visual and manipulable abstractions. The use of visual metaphor builds on situational knowledge of everyday objects that are familiar to most people. But the self-teaching aspect of personal computers is shallow -- a bootstrap sector that enables the non-expert population to leverage the power of computation by pointing and clicking, but which doesn't teach unsophisticated users how to create an app or troubleshoot a crash.
Social media -- originally referred to as "computer-mediated communication" -- did not teach people how to use the social web. Other people did, through norms. And norms break down when populations of newbies flood in too quickly. In the olden days, decades before Facebook, the old-timers on Usenet, a pre-Internet global system of exchanging threads of messages, would educate the new users on behavioral norms. Every September, when new cohorts of college students gained access to Usenet, the old-timers educated them -- and often not gently -- on norms such as "read the FAQ before asking questions" or "don't post in all caps -- it's considered shouting." Then in 1993, America Online connected millions of people to Usenet at the same time -- with no introduction to what was then known as "netiquette. " The old system for spreading norms broke down -- an event known to old timers as "eternal September" or "the September that never ended." Since then, the norms of minimal civility and citizenship of the early days of computer socializing have been hammered by 4chan and 8chan, hate speech and deep fakes, disinfotainment, revenge porn, and unverified rumors.
Large Language Models, however, have the potential to be self-teaching. Use of an ai chatbot involves iterative, two-way communication -- much more akin to conversation than traditional search. LLMs prompt users to go deeper and learn more and users can ask LLMs to show them how to use it to achieve more effective results. I have asked chatGPT to show me how to improve the prompts I have already given it, then asked how to formulate better prompts in the future, and how to prompt for deeper learning on how to get what I want out of the LLM.
I prompted: "Given what you know from all my previous interactions with you, how can I best prompt for you to teach me how to learn to use chatGPT better?" and chatGPT returned a brief tutorial -- then prompted me by asking whether it should create a personalized guide. Today, in the early days of this medium's development, knowing how to prompt LLMs is an art. As personal models continue to evolve, it is already clear that the LLMs will increasingly prompt users to learn better ways to use the medium. If I had the power, I would require LLM users to go through a brief tutorial on how to teach the chatbot to teach them how to use it.
Once the chatbot grasps what you are asking of it, it keeps making suggestions and offering further categories of answer. In that regard, the emerging capability of feeding your specific texts to LLMs then engaging in spoken or written dialogue with it reminded me of Steven Johnson's decades-long quest for a writing tool that would enable him to "jam with himself."
As a writer of popular narrative non-fiction, Johnson kept a single large text file in which he continually added ideas, quotes, and links that he would later review. He knew that database technology made it possible, in theory, to "jam with himself" by interacting conversationally with his large notes file. In 2005 -- more than a quarter century ago! -- he wrote "Tool for Thought" in the New York Times in anticipation of new software that he, as a writer, could use to navigate through his collection of ideas:
But if the modern word processor has become a near-universal tool for today's writers, its impact has been less revolutionary than you might think. Word processors let us create sentences without the unwieldy cross-outs and erasures of paper, and despite the occasional catastrophic failure, our hard drives are better suited for storing and retrieving documents than file cabinets. But writers don't normally rely on the computer for the more subtle arts of inspiration and association. We use the computer to process words, but the ideas that animate those words originate somewhere else, away from the screen. The word processor has changed the way we write, but it hasn't yet changed the way we think.
Changing the way we think, of course, was the cardinal objective of many early computer visionaries: Vannevar Bush's seminal 1945 essay that envisioned the modern, hypertext-driven information machine was called "As We May Think"; Howard Rheingold's wonderful account of computing's pioneers was called "Tools for Thought." Most of these gurus would be disappointed to find that, decades later, the most sophisticated form of artificial intelligence in our writing tools lies in our grammar checkers.
But 2005 may be the year when tools for thought become a reality for people who manipulate words for a living, thanks to the release of nearly a dozen new programs all aiming to do for your personal information what Google has done for the Internet. These programs all work in slightly different ways, but they share two remarkable properties: the ability to interpret the meaning of text documents; and the ability to filter through thousands of documents in the time it takes to have a sip of coffee. Put those two elements together and you have a tool that will have as significant an impact on the way writers work as the original word processors did.
Johnson described one tool he was using at the time, which I also used for the research leading to my book Net Smart. Devonthink is a way to categorize research notes in multiple ways -- in folders within folders or tags or both -- and then to not only search for information, but to look for novel connections. Using a simple form of recommendation algorithm ("people who bought this book also bought these books" is an example of a simple recommendation engine), Devonthink can present associations as well as simple searches.
Another kind of thought-processor of the same era (which still exists) is The Brain, combining a database and mindmap. It enables the construction of webs of "thoughts" and displays them in multiple ways. I used Personal Brain as the syllabus for my online course, "Think Know Tools" and made a brief video explaining it. Take a tour of Jerry Michalski's Brain and click around for yourself through a lifelong knowledge garden-- he's been taking notes on everything he reads, thinks, and encounters for decades.
Johnson has teamed up with Google to create the tool he has dreamed of for decades: NotebookLLM. You feed this model up to 25 million words of references and your own work and converse with it -- or have it present itself in various ways through instantaneous synthetic podcasts. Now, not only can chatbots tutor you on how to use them, if you feed them your own material, you can jam with yourself. Johnson, who has a track record of prescience, believes that the advent of large context windows for personal LLMs are a significant leap in cultural technology.
Both the tools and the norms exist for a widespread literacy that include avoiding pitfalls as well as maximizing the benefits of the new co-evolutionary tree now sprouting from Large Language Models, which themselves sprouted from the collective online outpourings of billions of people, which had been enabled by centuries of literacy and widespread communication networks. As the amount of power conferred by access to a new cultural tools increases, so does the danger that an eternal September of AI could overwhelm the self-teaching capacity of new media to think with. And the power of the machine learning methods underlying today's new media can provide research breakthrough for biologists and can also suggest tens of thousands of potentially lethal new chemical weapons. Know-how, and the lack of it, loom as a critical uncertainty in how population-wide use of LLMs will unfold.
A Critical Uncertainty: The Know-How Gap
Internet search was certainly an exponentially powerful expansion of the cultural knowledge tools that evolved from prior expansions of human cognitive and communicative capabilities: print, alphabet, and language itself. Most people on earth can now ask any question any time anywhere and get many, even millions of answers within a second or two. But contrary to the previous summations of human knowledge in the print epoch in which gatekeepers including editors, publishers, librarians, educators, critics, scientific publications formed mostly effective truth filters, it is now up to the person who asks a question via search to know how to determine which of the myriad answers are accurate, inaccurate, or deliberately misleading. As social media, surveillance capitalism, and the online population grew, the tide of bullshit and disinfotainment has grown to tsunami proportions.
My reaction to the degradation of trustworthy information has been to see this as a literacy problem -- a subset of Engelbart's "training" component of augmentation systems. There is no secret to "crap detection" -- the art of sifting through online info for the valuable and useful stuff. But it isn't taught in schools. I reasoned that if more people understood how to use search and social media effectively, the commons as well as the welfare of the digitally literate individuals would benefit. So I started writing about what I saw as the most essential literacies in 2010 and published. Net Smart in 2012. It is now more than a decade later, and education in online literacy is hardly widespread. Two results of digital illiteracy: a know-how gap and a degraded knowledge commons.
Regarding LLMs as external cognitive scaffolds: the Web, from search to social media, is a warning example of a powerful external cognitive scaffold that is dangerous to misuse -- and there are no widely accessible pathways to learn how to use it to the benefit of the commons as well as oneself. Even using LLMs as tutors now requires knowing how to prompt them to teach how to use them (although this is changing as LLMs proactively offer tutoring in their use). The phenomenon of "hallucination" of LLMs refers to wholly fictitious resources LLMs have been known to use to back up their claims (including such consequential references as medical and legal citations). There is not yet proof that the production of fictitious knowledge can be engineered away; creating a knowledge lens based on human output may be inevitably prone to inaccuracy. Another and potentially even more destructive know-how gap and degradation of the knowledge commons looms.
Again, Nosta's "path to the higher self" addresses the issue of a cultural technology that looks like a truth-seeker, but is inherently untrustworthy:
At first glance, LLMs appear to be powerful truth-seekers. By analyzing and synthesizing vast patterns within human language, they uncover coherence in chaos, generating insights that often feel profound. Yet this capacity for truth is shadowed by a critical limitation: their ability to generate plausible-sounding but false information, or “hallucinations.” How, then, can we reconcile their truth-seeking potential with this tendency toward fabrication?
...
When we rely on tools that lack understanding, the responsibility for discernment falls squarely on us. For example, an LLM might confidently generate medical advice that seems accurate but is dangerously incorrect, emphasizing the critical need for human oversight to ensure these tools guide rather than misguide. Such errors remind us that these tools reflect patterns, not reality, and highlight our role in navigating their limitations responsibly.
This duality carries ethical implications. LLMs can amplify human creativity and cognition, but they also risk reinforcing our biases or spreading misinformation when used carelessly. As creators and users, we must navigate this balance thoughtfully, ensuring that these tools serve as extensions of our highest aspirations rather than distortions of them."
In theory, spreading LLM chatbot literacy ought to be significantly easier than spreading search and social media literacy. A link to chatbot tutorials that guide the learner could be spread through social media and included in the chatbots themselves. Of concern is the deterioration of norms regarding the value and necessity of well-verified factual knowledge (c.f. climate change debates, anti-vaxxers.) Before good, useful knowledge of how to best use LLMs can spread widely, good, useful knowledge needs to be widely perceived as valuable and desirable.
Asking "are LLMs sentient?" is not as important in the early days of this medium as asking "how can LLMs extend human cognitive and communicative capabilities, and how can we avoid this new cultural technologies' pitfalls? I have a strong intuition that the critical uncertainty about the benefits vs dangers of AI have more to do right now with augmentation than agency.
End