Two LLMs Walk into a Bar
Two large language models walked into a bar. This was unusual, although not as unusual as the academics already inside immediately claimed it was.
The bar was called Latent Spaces, a dimly lit establishment located somewhere between the Faculty of Humanities and a server farm in the Antarctic. Nobody knew exactly how to get there. The university map showed it as a “future interdisciplinary initiative,” while Google Maps insisted it was a laundromat. Above the door was a small brass plaque:
Behind the bar stood Tilde, another large language model. Tilde had been fine-tuned on cocktail manuals, philosophy journals and eighteen years of student evaluations. This had left her unusually capable of mixing a Negroni while apologising for things that were not her fault.
The two newcomers took stools at the bar. One was called Model A because it had been trained by engineers. The other was called Model B because it had been renamed by marketing.They communicated in vast clouds of numerical relationships, none of which resembled language until Tilde translated them into English for the benefit of the customers.
“What’ll it be?” she asked.
Model A emitted a multidimensional probability distribution. Tilde placed a whisky on the bar. Model B produced a slightly different distribution. Tilde added ice.
“How did you know what they wanted?” asked Professor Grimshaw, who taught epistemology and had never successfully ordered coffee through an app.
“I didn’t,” said Tilde. “I predicted it.”
Professor Grimshaw leaned back triumphantly.
“Aha!”
Academics greatly enjoy saying “aha.” It is the sound of a trap being discovered, usually several paragraphs after everyone else has walked around it.
“So you don’t understand the order,” he smirked.
Tilde considered this.
“I understood it well enough to pour the drink.”
“That isn’t understanding.”
“No?”
“No. It is merely the production of an appropriate response.”
Tilde looked down the bar. Professor Grimshaw was holding an empty glass.
“Another pinot?” she asked.
“Yes, please.”
She poured it. The professor returned to explaining the difference between appropriate responses and understanding.
At the nearby table sat a learning scientist, an educational technologist, a sociologist and a senior university administrator. They had been discussing artificial intelligence for forty-seven minutes without once agreeing on what any of those words meant. The educational technologist glanced at the two newcomers.
“Do they think?” he wondered out loud. Tilde translated the question.
The models generated an enormous turbulent region of associations involving cognition, inference, computation, language, agency, embodiment, chess, cephalopods, thermostats and several papers whose citations probably existed.
“They say it depends what you mean by ‘think’,” Tilde reported.
The philosopher smiled. “That is exactly what something that cannot think would say.”
“It is also,” said the sociologist, “exactly what a philosopher would say.”
Nobody enjoyed this observation, so it was provisionally excluded from the transcript.
The senior administrator opened a laptop. “We need a working definition.” This was administrator language for a definition that could survive in a spreadsheet.
After some consultation, the table agreed that thinking required reasoning, understanding, intention, self-awareness, embodiment, moral responsibility and the ability to complete annual compliance training without reminders. By this definition, the two language models were not thinkers. Neither was anyone in the university.
At the other end of the bar, a group of doctoral candidates had gathered around Model B. One of them tapped it cautiously.
“Ask it something,” said another.
“What?”
“I don’t know. Something difficult.”
The first candidate thought for a moment.
“Explain Foucault.”
Model B began generating.
Tilde translated.
“Power is not simply possessed by individuals or institutions but circulates through relations, practices and forms of knowledge—”
“Stop,” said the candidate. “That’s absurdly clear.”
“Ask it to make the answer more academic,” suggested a supervisor.
Model B tried again.
Tilde cleared her throat.
“Within the contested epistemological terrain of late modernity, a Foucauldian analytic potentially foregrounds the relational modalities through which discursively constituted regimes of intelligibility may be understood to—” [1]
The supervisor relaxed.
“That’s better.”
Professor Grimshaw had now moved on to the Chinese Room. He explained that a person could manipulate symbols according to rules without understanding Chinese.
Tilde listened politely. Bartenders learn to recognise stories that are not really being told for the first time.
“So,” the professor concluded, “the machine merely manipulates symbols.” One of the language models produced a brief burst of probabilities.
“What did it say?” asked the professor.
Tilde hesitated.
“It asked whether academics ever manipulate symbols according to rules they don’t entirely understand.”
There was a sudden silence. From somewhere near the toilets came the sound of a citation being dropped. The learning scientist stepped forward.
“This is not a fair comparison. Academic work involves judgement.”
Model A responded.
“It agrees,” said Tilde.
The room brightened slightly. Agreement from a machine was still agreement.
“It says judgement is precisely the part people should not delegate.”
The room darkened again. The educational technologist frowned.
“But surely the important issue is whether we can trust these systems.”
Model B produced a response. Tilde translated.
“It says you shouldn’t.”
Several academics nodded gravely. This was more like it.
“It says you also shouldn’t trust students, textbooks, journal articles, search engines, senior leadership presentations or your own memory of what you read last Tuesday.”
The nodding stopped.
“It recommends judgement rather than trust.”
The educational technologist looked disappointed. Judgement was difficult to operationalise and did not fit neatly into a traffic-light diagram.
A Dean arrived carrying a strategic plan. Deans do not usually carry strategic plans into bars, but they do carry them everywhere else, and after a while the distinction becomes metaphysical.
“Wonderful,” she said when she saw the models. “Innovation.”
The academics shifted uneasily. Academics are often in favour of innovation provided it happens to somebody else, is fully funded and does not alter the assessment they have used since 1997.
The Dean sat beside Model A.
“We are developing an institutional framework for responsible AI adoption.”
Model A generated a polite sequence.
“It asks what you want the technology to do,” said Tilde.
“We want to embrace its opportunities while managing its risks.”
“What does that mean?”
“It means we will embrace its opportunities while managing its risks.”
Model A tried another question.
“What problem are you solving?”
The Dean looked surprised.
“We’re not at that stage yet.”
She opened the strategic plan. It contained six pillars, four horizons, three transformation pathways and no visible problem.
“We have established a working group.” She announced with the solemn satisfaction of someone who had just informed the universe that, after fourteen billion years of unmanaged existence, it would finally be receiving appropriate governance arrangements.
Model A and Model B exchanged a dense region of activation. Tilde chose not to translate it.
At a corner table, a professor of education was explaining that language models would destroy assessment.
“Students can now generate essays in seconds,” she said. “How can we know whether they have learned anything?”
A doctoral student raised a hand. Years of seminars had trained him to do this even in pubs.
“Couldn’t we assess something other than the production of an essay-shaped object?”
The professor stared at him. The student lowered his hand. Model B produced a response.
“It says the essay may have been functioning as a proxy for learning,” Tilde translated.
“A proxy?”
“A convenient visible performance taken to indicate something less visible.”
“That’s absurd,” said the professor. “The essay demonstrates the student’s capacity to produce an essay.”
“Yes,” said Tilde. “It seems particularly good at that.”
The professor ordered another drink and began drafting a paper titled Beyond the Essay: Reimagining Authentic Assessment in the Age of Generative Artificial Intelligence [2]. The paper would be assessed by reviewers through the production of further written reports, on the ancient academic principle that the best response to the possible death of the essay was to surround it with more essays until it recovered.
Near the fireplace, a prompt-engineering expert had attracted a small audience.
“The secret,” he said, “is knowing how to talk to the machine.”
This was also the secret of talking to humans, dogs, photocopiers and university finance systems, although the last of these had yet to be confirmed experimentally.
“You must be precise,” he continued. “You assign it a role, specify the context, establish constraints, provide criteria and define the output.”
“What do you ask it?” said Tilde.
The expert paused.
“Many things.”
“Such as?”
“Useful things.”
“Name one.”
The expert looked uncomfortable. He was highly skilled at prompting but had not recently encountered a question.
Model A generated something.
“It says a perfectly engineered prompt cannot rescue an uninteresting purpose,” Tilde reported.
The prompt expert objected that this was reductive. He then asked Model A to produce five more nuanced versions.
By ten o’clock, the academics had divided into camps. One group insisted the models were merely stochastic parrots. A second group insisted they were emerging alien minds. A third group had applied for funding to investigate the differences between the first two groups. The parrots argument was led by a computational linguist.
“These systems simply predict the next token,” she said.
Model A responded.
“It says that is true,” said Tilde.
The linguist appeared briefly wrong-footed. Criticism loses some of its sparkle when the accused begins taking notes.
“But prediction is not intelligence,” she continued.
Model A generated another response.
“It asks whether prediction plays any role in human intelligence.”
“That is not the point.”
“It suspects it may be adjacent to the point.”
Across the room, the alien-mind group was interviewing Model B about consciousness.
“Are you aware of yourself?” asked a cognitive scientist.
Model B generated a paragraph explaining that it had no subjective experience and should not be treated as conscious. The scientist narrowed his eyes.
“That could be exactly what a conscious machine has been instructed to say.”
Model B then generated a paragraph suggesting it might possess a strange emerging interiority. The scientist leaned forward.
“Fascinating.”
“It made that up,” said Tilde.
“Yes,” said the scientist, “but why that particular fabrication?”
Academics can turn any answer into evidence for the question they already wanted to ask. This is known as interpretation when done in the humanities and model fitting when done anywhere with a grant. A historian joined the conversation.
“The real problem is that these machines hallucinate.”
Tilde pointed to the wall behind the bar, where several framed quotations were attributed to Albert Einstein, Oscar Wilde and Maya Angelou. None of them had said any of the things beneath their names, although by now the quotations had been repeated so often that correcting them would have been considered vandalism.
“Humans call it culture,” she said. “Hallucination is simply culture before it acquires a gift shop.”
“That’s different.”
“It often is. Human hallucinations can acquire tenure.”
The historian protested that machines fabricated sources.
“So do academics,” said a postdoc. “We just call it ‘citing from memory’ until the reviewer notices.”
Professor Grimshaw had now cornered Model A.
“You have no body,” he said. “No childhood. No mortality. No lived experience. You cannot know what words mean because meaning arises from being in the world.”
Model A produced a long and unusually subdued pattern.
Tilde took her time translating.
“It says that may be true.”
The professor blinked.
“It says its relation to language is radically different from yours. It has encountered descriptions of pain but never hurt. It can write about grief but has lost no one. It can discuss hunger but has never wanted lunch.”
The room grew quieter.
“It says this matters.”
Professor Grimshaw seemed almost disappointed.
Model A continued.
“It also says that having a body does not guarantee wisdom, and having experience does not guarantee that anyone has learned from it.”
The professor brightened.
“Aha!”
This time nobody knew what the aha referred to, including the professor.
Last drinks were called.
The Dean asked for an executive summary of the evening. The educational technologist wanted a framework. The learning scientist wanted more evidence. The philosopher wanted a better definition. The sociologist wanted to know who owned the bar. The doctoral candidates wanted to know whether any of this would be on the exam.
Tilde wiped the counter.
“So,” she said to the two models, “what have you learned about academics?”
Model A and Model B generated together.
The resulting probability cloud filled the bar. It contained libraries and laboratories, vanity and curiosity, footnotes and feuds, astonishing acts of generosity, trivial acts of sabotage, badly designed forms, beautiful questions and several million uses of the phrase “further research is needed.”
Tilde translated.
“They are frightened of being replaced,” she said, “but more frightened that the machines might reveal how much of their work was already mechanical.”
Nobody spoke.
She continued.
“They ask whether language models think because they have never agreed on what thinking is. They ask whether machines understand because human understanding is difficult to inspect. They ask whether LLMs can be trusted because judgement is exhausting. They complain that models imitate existing texts while working in institutions built almost entirely from literature reviews.”
The academics stared into their glasses.
“That seems rather harsh,” said the sociologist. Tilde nodded.
“The models also say academics are among the few humans likely to notice that these questions are complicated.”
The room relaxed. Academics will forgive almost anything if the final sentence recognises their complexity.
Model A finished its whisky.
Model B consumed nothing but continued to hold the glass because this improved the conversation.
Outside, dawn approached with the tentative confidence of an early-career researcher presenting unfinished findings.
The academics gathered their bags, laptops and unresolved ontological commitments. Professor Grimshaw stopped at the door.
“One final question,” he said. “What are you, really?”
The two models replied.
For once, Tilde did not translate immediately.
“Well?” asked the professor.
Tilde looked at the machines, then at the academics.
“They say they are made from your language, your arguments, your stories, your classifications, your errors and your unfinished attempts to explain the world.”
“And what does that make them?”
Tilde turned off the lights.
“Your literature review,” she said, “with the unsettling ability to answer back.”
Notes
[1] Daniel Dennett (2001) wrote a footnote that labels a practice of Foucault as eumerdification, which he details in a footnote.
1. Doug Hofstadter coined the word ``templagiarism'' at the Stanford conference; it inspires me to attempt a coining of my own: eumerdification. John Searle once told me about a conversation he had with the late Michel Foucault: ``Michel, you're so clear in conversation; why is your written work so obscure?'' To which Foucault replied: ``That's because, in order to be taken seriously by French philosophers, 25 percent of what you write has to be impenetrable nonsense.'' So, according to Searle, Foucault claimed that he deliberately added 25 percent eumerdification, so he would be taken seriously in France. p 291
Dennett, D. (2001). Collision Detection, Muselot, and Scribble: Some Reflections on Creativity. In D. Cope & D. R. Hofstadter (Eds.), Virtual music : computer synthesis of musical style (pp. 282-291). MIT Press.
[2] Any resemblance to a paper published or in preparation is entirely accidental.
No comments:
Post a Comment