The Question Machine
The arrival of large language models created an educational emergency. This was fortunate, because universities are extremely good at educational emergencies. Within months there were task forces, principles, frameworks, guidelines and webinars. Documents appeared containing diagrams with arrows in them. Somewhere, almost certainly, a committee was established to coordinate the work of the other committees.
The problem seemed obvious. An Answer Machine had been invented. Students could type in a question and receive an answer. Worse, the answer arrived quickly, confidently and without requiring them to find a library book, remember a password or sit through a PowerPoint containing the words learning outcomes. Clearly, something had gone terribly wrong.
Tim Klapdor, in a recent post called The Answer Machine, makes the larger point that humans have always been attracted to things that promise answers. Religions have done it. Experts have done it. Universities have done it. Search engines have done it. Now we have machines that will cheerfully answer almost anything, including questions nobody particularly needed answering. But I wonder if we are giving the answer rather too much credit.
For most of the history of formal education, answers were expensive. They lived in books, libraries, laboratories, disciplines and the heads of people who had spent twenty years learning when to say it depends.
Education built an impressive amount of machinery around them. Curricula specified answers. Teachers transmitted them. Textbooks stored them. Examinations asked students to reproduce them. Universities certified that people possessed an appropriate quantity of them.
But then answers became cheap. Not necessarily good. Not necessarily true. Not necessarily useful. Just ridiculously cheap at the point of use — provided you did not look too closely at what it costs to make them cheap. The bill, as usual, had been sent somewhere else.
You can now obtain 800 fluent, plausible words on almost anything before you have had time to regret asking for them. This is rather more interesting than whether students might use ChatGPT to write an essay. If answers have become cheap, perhaps answers are no longer where the interesting educational work is. Perhaps the scarce thing is the question.
Not the elaborately engineered prompt beloved of the AI productivity industry:
Act as an internationally recognised expert in medieval drainage systems and produce a seven-point framework...
I’m thinking about the question that opens up something new. The output that makes you notice what you had not noticed, or puzzle over what had previously seemed unremarkable. The useful answer is not an endpoint but an opening into a Kauffmanesque adjacent possible: a new patch of intellectual territory from which more interesting questions suddenly become askable. The answer now matters because it gets you to a place from which you can ask a better question.”The question you could not have asked before receiving the previous answer.
Seen this way, a LLM starts looking less like an Answer Machine and more like a rather odd looking Question Machine.
You ask it something. It replies. The reply maybe is wrong, banal, surprising, interesting or, occasionally, much better than you expected. That changes what you know. That changes what you now may want to ask. So you ask something else. And now you are somewhere you could not have been when you began prompting. The interesting question is no longer:
Did the machine give the correct answer?
It is:
Where did this exchange take me?
This makes the institutional responses to LLM use look somewhat odd. Something genuinely strange arrives and universities immediately ask:
Is it cheating? What is acceptable use? How do we detect it? What policy should we adopt? Can someone develop a framework by Tuesday?
These are not bad questions. They are questions with a strong desire to stop being questions. Something totally unfamiliar appears over the hill and the institutional machinery starts shouting, Darlek like:
DOMESTICATE! DOMESTICATE!
The thing must be captured, classified and placed in a policy document, preferably one with numbered headings and a tasteful diagram showing Responsible AI Use in the middle.
But perhaps this rush to domesticate is precisely the wrong instinct. We have barely begun to see what happens when people think, write, argue, design, learn and become productively confused in the presence of these machines. It seems a little early to put them on a lead, give them a policy number and declare the experiment complete.
Perhaps we need fewer confident declarations about the impact of AI on education and rather more small-r research. Yeah. I have a thing for small-r research.
Small-r research begins with the deeply unfashionable sentence:
I don't know. Let's try something.
What happens if I tell the machine to argue with me rather than applaud? What happens if the student looks at its beautifully polished suggestion and says, “Nope”? Why did yesterday’s prompt open a door and today’s produce beige intellectual porridge? What happens if I give the machine a different job entirely — critic, provocateur, idiot companion, unreliable witness? What happens if I leave it out of the room? And why can two people sit down with exactly the same model and emerge having visited completely different intellectual planets? This is a more poke the beast and see what it does approach.
Clearly, these questions are unlikely to produce a national framework. This is part of their charm. They amount to poking the thing with a stick and paying attention.
The standard picture of LLM use is: human asks then machine answers
That seems to leave out the interesting bit.
A human asks a question. The machine produces something. The human reads it. The human is now, however slightly, a different human as a consequence of what the machine spat out. The next question therefore can come from somewhere new. The useful object of interest may not be the prompt. It may not be the answer. It may be the trajectory.
This is why I am suspicious of tidy distinctions between machine answers and human wisdom.
It is wonderfully reassuring to put the machine over there, extruding synthetic sludge by the bucketful, while we humans remain over here being deep, relational, embodied and, on a good day, wise. Unfortunately, humans have also produced committee minutes, airport novels, management jargon and several centuries of confidently wrong ideas. Reality is seldom kind enough to respect the categories we invent for it.
A conversation with a LLM might send me back to a book, into an argument with a colleague, towards an experiment, or, more alarmingly, into the discovery that something I have confidently believed for twenty years, a load-bearing chunk of my intellectual path dependence, is held together by habit, professional muscle memory and a small republic of papers citing one another in a reassuring circle.
I think it’s unhelpful to want the machine to be wise. It merely has to perturb me. Sometimes productively. Sometimes disastrously. Sometimes by inventing three Belgian researchers who have never existed. Which is precisely where judgement becomes interesting. If answers become abundant, judgement becomes way more important, not less.
Am I able to recognise an interesting wrong answer? Do I notice when the machine has simply polished my assumptions and handed them back? Can I tell when something deserves pursuing? Do I ignore something merely because it sounds authoritative? Does the output nudge me to ask the next question?
Those seem to me rather more demanding capacities than producing a five-paragraph essay on the causes of the First World War. Although I would be interested to know how much geopolitical carnage can now be compressed into five paragraphs without violating the rubric. Or perhaps the real breakthrough is discovering that the Schlieffen Plan was, in fact, paragraph three.
Perhaps this is the educational opportunity hidden inside the educational emergency. We built institutions for a world in which answers were scarce. Now we have machines producing them in industrial quantities. The domesticating response is to defend the old scarcity. Ban the machine. Restrict it. Detect it. Require students to prove that the answer came from somewhere sufficiently inconvenient.
The other possibility is more unsettling. We could ask what education becomes when the answer is no longer the star of the show. Which things are still worth learning? What deserves assessment? What should we get better at noticing? And, perhaps most awkwardly for institutions built around approved answers, which questions should we stop trying to house-train?”
We may eventually discover that LLMs are disastrous for education. We may discover they are transformative. More likely, both statements will turn out to be annoyingly true, often in the same classroom before lunch.
But for now, perhaps we could resist the urge to issue a verdict. The history of AI prediction is not exactly a monument to human foresight. We have been confidently announcing both the imminent arrival and imminent failure of artificial intelligence for decades, usually shortly before being surprised by something else entirely.
Universities have spent centuries turning strange things into familiar answers. Then a machine arrived that could manufacture familiar answers in twelve seconds and everyone reached for the emergency procedures manual.
Perhaps we have been protecting the wrong end of the business. Education should also make the familiar strange: unsettle the obvious, annoy the settled, and occasionally discover that the intellectual furniture has been nailed to the floor for no particularly good reason.
So perhaps the interesting question is not what the machine knows, but what becomes possible to ask once answers are cheap. This is, of course, far too important to be left to curiosity. A working group will now determine which questions are permissible.