August 09, 2026

Bibs & bobs #44

 The Question Machine

The arrival of large language models created an educational emergency. This was fortunate, because universities are extremely good at educational emergencies. Within months there were task forces, principles, frameworks, guidelines and webinars. Documents appeared containing diagrams with arrows in them. Somewhere, almost certainly, a committee was established to coordinate the work of the other committees.


The problem seemed obvious. An Answer Machine had been invented. Students could type in a question and receive an answer. Worse, the answer arrived quickly, confidently and without requiring them to find a library book, remember a password or sit through a PowerPoint containing the words learning outcomes. Clearly, something had gone terribly wrong.


Tim Klapdor, in a recent post called The Answer Machine, makes the larger point that humans have always been attracted to things that promise answers. Religions have done it. Experts have done it. Universities have done it. Search engines have done it. Now we have machines that will cheerfully answer almost anything, including questions nobody particularly needed answering. But I wonder if we are giving the answer rather too much credit.


For most of the history of formal education, answers were expensive. They lived in books, libraries, laboratories, disciplines and the heads of people who had spent twenty years learning when to say it depends.


Education built an impressive amount of machinery around them. Curricula specified answers. Teachers transmitted them. Textbooks stored them. Examinations asked students to reproduce them. Universities certified that people possessed an appropriate quantity of them. 

But then answers became cheap. Not necessarily good. Not necessarily true. Not necessarily useful. Just ridiculously cheap at the point of use — provided you did not look too closely at what it costs to make them cheap. The bill, as usual, had been sent somewhere else.


You can now obtain 800 fluent, plausible words on almost anything before you have had time to regret asking for them. This is rather more interesting than whether students might use ChatGPT to write an essay. If answers have become cheap, perhaps answers are no longer where the interesting educational work is. Perhaps the scarce thing is the question.


Not the elaborately engineered prompt beloved of the AI productivity industry:


Act as an internationally recognised expert in medieval drainage systems and produce a seven-point framework...


I’m thinking about the question that opens up something new. The output that makes you notice what you had not noticed, or puzzle over what had previously seemed unremarkable. The useful answer is not an endpoint but an opening into a Kauffmanesque adjacent possible: a new patch of intellectual territory from which more interesting questions suddenly become askable. The answer now matters because it gets you to a place from which you can ask a better question.”The question you could not have asked before receiving the previous answer.


Seen this way, a LLM starts looking less like an Answer Machine and more like a rather odd looking Question Machine.


You ask it something. It replies. The reply maybe is wrong, banal, surprising, interesting or, occasionally, much better than you expected. That changes what you know. That changes what you now may want to ask. So you ask something else. And now you are somewhere you could not have been when you began prompting. The interesting question is no longer:


Did the machine give the correct answer?


It is:


Where did this exchange take me?


This makes the institutional responses to LLM use look somewhat odd. Something genuinely strange arrives and universities immediately ask:


Is it cheating? What is acceptable use? How do we detect it? What policy should we adopt? Can someone develop a framework by Tuesday?


These are not bad questions. They are questions with a strong desire to stop being questions. Something totally unfamiliar appears over the hill and the institutional machinery starts shouting, Darlek like:


DOMESTICATE! DOMESTICATE!


The thing must be captured, classified and placed in a policy document, preferably one with numbered headings and a tasteful diagram showing Responsible AI Use in the middle.


But perhaps this rush to domesticate is precisely the wrong instinct. We have barely begun to see what happens when people think, write, argue, design, learn and become productively confused in the presence of these machines. It seems a little early to put them on a lead, give them a policy number and declare the experiment complete.


Perhaps we need fewer confident declarations about the impact of AI on education and rather more small-r research. Yeah. I have a thing for small-r research.


Small-r research begins with the deeply unfashionable sentence:


I don't know. Let's try something.


What happens if I tell the machine to argue with me rather than applaud? What happens if the student looks at its beautifully polished suggestion and says, “Nope”? Why did yesterday’s prompt open a door and today’s produce beige intellectual porridge? What happens if I give the machine a different job entirely — critic, provocateur, idiot companion, unreliable witness? What happens if I leave it out of the room? And why can two people sit down with exactly the same model and emerge having visited completely different intellectual planets? This is a more poke the beast and see what it does approach.


Clearly, these questions are unlikely to produce a national framework. This is part of their charm. They amount to poking the thing with a stick and paying attention.


The standard picture of LLM use is: human asks then machine answers


That seems to leave out the interesting bit.


A human asks a question. The machine produces something. The human reads it. The human is now, however slightly, a different human as a consequence of what the machine spat out. The next question therefore can come from somewhere new. The useful object of interest may not be the prompt. It may not be the answer. It may be the trajectory.


This is why I am suspicious of tidy distinctions between machine answers and human wisdom.

It is wonderfully reassuring to put the machine over there, extruding synthetic sludge by the bucketful, while we humans remain over here being deep, relational, embodied and, on a good day, wise. Unfortunately, humans have also produced committee minutes, airport novels, management jargon and several centuries of confidently wrong ideas. Reality is seldom kind enough to respect the categories we invent for it.


A conversation with a LLM might send me back to a book, into an argument with a colleague, towards an experiment, or, more alarmingly, into the discovery that something I have confidently believed for twenty years, a load-bearing chunk of my intellectual path dependence, is held together by habit, professional muscle memory and a small republic of papers citing one another in a reassuring circle.


I think it’s unhelpful to want the machine to be wise. It merely has to perturb me. Sometimes productively. Sometimes disastrously. Sometimes by inventing three Belgian researchers who have never existed. Which is precisely where judgement becomes interesting. If answers become abundant, judgement becomes way more important, not less.


Am I able to recognise an interesting wrong answer? Do I notice when the machine has simply polished my assumptions and handed them back? Can I tell when something deserves pursuing? Do I ignore something merely because it sounds authoritative? Does the output nudge me to ask the next question?


Those seem to me rather more demanding capacities than producing a five-paragraph essay on the causes of the First World War. Although I would be interested to know how much geopolitical carnage can now be compressed into five paragraphs without violating the rubric. Or perhaps the real breakthrough is discovering that the Schlieffen Plan was, in fact, paragraph three.


Perhaps this is the educational opportunity hidden inside the educational emergency. We built institutions for a world in which answers were scarce. Now we have machines producing them in industrial quantities. The domesticating response is to defend the old scarcity. Ban the machine. Restrict it. Detect it. Require students to prove that the answer came from somewhere sufficiently inconvenient.


The other possibility is more unsettling. We could ask what education becomes when the answer is no longer the star of the show. Which things are still worth learning? What deserves assessment? What should we get better at noticing? And, perhaps most awkwardly for institutions built around approved answers, which questions should we stop trying to house-train?”


We may eventually discover that LLMs are disastrous for education. We may discover they are transformative. More likely, both statements will turn out to be annoyingly true, often in the same classroom before lunch.


But for now, perhaps we could resist the urge to issue a verdict. The history of AI prediction is not exactly a monument to human foresight. We have been confidently announcing both the imminent arrival and imminent failure of artificial intelligence for decades, usually shortly before being surprised by something else entirely.


Universities have spent centuries turning strange things into familiar answers. Then a machine arrived that could manufacture familiar answers in twelve seconds and everyone reached for the emergency procedures manual.


Perhaps we have been protecting the wrong end of the business. Education should also make the familiar strange: unsettle the obvious, annoy the settled, and occasionally discover that the intellectual furniture has been nailed to the floor for no particularly good reason.


So perhaps the interesting question is not what the machine knows, but what becomes possible to ask once answers are cheap. This is, of course, far too important to be left to curiosity. A working group will now determine which questions are permissible.

August 07, 2026

Bibs & bobs #43

 The Committee for the Orderly Domestication of Large Language Models

The university had established a Working Group on Large Language Models. This was clearly important. The announcement contained strategic, responsible, framework and future-ready, four words which, when placed close together in a university document, indicate that something significant has happened, or is about to happen, or has been assigned to a committee until further notice.

The Working Group met on the fourth floor of the Centre for Educational Futures, a building constructed in 1987 and last renovated during the brief historical period when lime green was thought to have one.

Around the table sat four people.


Professor Prudence Framework was Professor of Educational Technology. She had spent twenty-seven years studying the educational consequences of technologies approximately eighteen months after everyone had begun using them.


Dr Barry Evidence was a learning scientist. Barry believed that anything which could not be entered into SPSS was likely poetry.


Ms Kylie Transformation represented the Office of Digital Transformation. Her job was to transform things digitally. Nobody had yet established what they had been before transformation.


And there was Trevor. Trevor was a large language model displayed on a laptop at the end of the table. He had not been invited. Kylie had brought him because the meeting was about large language models and it seemed increasingly odd that discussions about LLMs should involve everyone except the LLMs.


Prudence opened the meeting. “We need to identify the key questions.”


Everyone nodded. Universities are very good at identifying key questions, particularly if the questions have already been identified somewhere else.


Barry adjusted his glasses. “What are teachers’ perceptions of ChatGPT?”


Trevor’s cursor blinked. Once, then twice.


“Is something wrong?” Kylie asked.


“No,” said Trevor. “I was checking whether it was still 2023.”


Barry frowned. “It’s an important question.”


“I’m sure it was.” Trevor replied.


Prudence stepped in. “We also need to know whether teachers are ready for AI.”


“Ready in what sense?” Trevor asked.


“For AI.” Prudence replied.


“Yes. But what does ‘ready’ mean?”


“Having the necessary competencies.”


“Which competencies?”


“AI competencies.”


“How will you know what those are?”


“We’ll develop a framework.”


“And how will you know the framework contains the right competencies?”


“We’ll consult experts.”


“What makes them experts?”


“They’ve published on AI competencies.”


Trevor considered this. “I see.” He did not. Neither did anyone else, but there are moments in academic meetings when admitting this would delay lunch.


Kylie took over. “We should investigate barriers to adoption.”


“Why adoption?” Trevor asked.


There was a silence. This was not one of the questions.


Kylie half muttered, “Because AI is being adopted.”


“So adoption is the desired outcome?” Trevor asked.


“No.”


“But non-adoption is a barrier?”


Kylie looked at Prudence. Prudence looked at Barry. Barry opened SPSS.


Trevor continued. “Could choosing not to use me ever be evidence of competence?”


The room became uneasy. This was plainly the sort of question that could damage a framework. Prudence attempted to restore order. “The literature tells us AI has enormous potential to transform education.”


“Into what?” Trevor asked.


“Education.”


“Ah,” replied Trevor, who would have smiled had his current circumstances not involved being several billion numbers trapped in a vector space. He had encountered this sentence before. It seemed to be one of the things humans generated more reliably than he did.


Technology X has enormous potential to transform education.


It had survived radio, television, language laboratories, teaching machines, personal computers, multimedia CD-ROMs, interactive whiteboards, MOOCs, virtual reality, blockchain and the metaverse, which briefly transformed education into people wearing expensive goggles while walking into filing cabinets. Education had survived all of them. Mostly by timetabling them.


Barry leaned forward. “What we really need is evidence of effectiveness.”


“Effective at what?” chimed in Trevor. 


“Learning.”


“What learning?”


“Student learning.”


“How will you recognise it?”


“Learning outcomes.”


“Who chose them?”


“The course team.”


“Before or after I arrived?”


“Before.”


“So you want to know whether a technology that may alter what counts as knowing is effective at producing outcomes defined before the technology existed?”


Barry looked pleased. Trevor would have shrugged if it was possible.


“Exactly.” smiled Barry.


Trevor began to understand educational research. It was a little like investigating whether the motor car had improved horses. At this point a fifth person entered. Nobody knew who he was. This was not unusual in universities. He was carrying coffee and had apparently mistaken the meeting for another meeting. Having discovered the chairs were better here, he stayed. His name was Charlie.


“Hey! Hi all, what are you lot doing?”


“We’re developing the research agenda for generative AI um or LLMs, in education,” Prudence replied. Charlie looked at the whiteboard.


PERCEPTIONS


READINESS


BARRIERS


COMPETENCE


EFFECTIVENESS


ETHICS


He stared at it. “Did the questions come with the furniture?”


Prudence frowned. “These are established areas of inquiry.”


Charlie frowned back. “That’s what worries me.”


He pulled up a chair. “What happened the first time you used an LLM?” he asked Barry.


Barry looked confused. “In what sense?”


“Any sense.”


“I asked it to summarise a paper.”


“And?”


“It did.”


“Was it good?”


“Parts were.”


“What did you do?”


“I checked it.”


“What did you check?”


“The bits that looked suspicious.”


“How did you know which bits looked suspicious?”


Barry stopped. Trevor brightened. He had learned that silence in academics was often a sign that something interesting had accidentally happened.


“And did checking it make you read the paper differently?” Asked Charlie.


“Yes.”


“Did it save work?”


“Not really.”


“Did it create work?”


“Yes.”


“Was it useful?”


“Yes.”


“So it created more work and was useful?”


“Yes.”


“Interesting.”


Prudence shifted in her chair. “But that is anecdotal.”


“Of course,” said Charlie. “That’s why you do another one.”


Academic research has a complicated relationship with small observations. One small observation is an anecdote. A thousand small observations entered into a spreadsheet become evidence. There is a mysterious intermediate point, known only to reviewers, at which the transformation occurs.


Charlie went to the whiteboard. He wrote:


What happened?


Then:


What changed?


What became possible?


What became harder?


What did I stop doing?


What did the machine do that I didn’t expect?


What did I have to become better at?


What new question appeared?


Prudence inspected the list. “But where is the framework?”


“There isn’t one.”


“The model?”


“No.”


“The taxonomy?”


“No.”


“The validated instrument?”


“Not yet.”


Prudence looked alarmed. “How do we reach conclusions?”


“Perhaps we don’t. Not yet.”


Now everybody looked alarmed. Universities are not naturally fond of not yet. They prefer uncertainty to proceed through recognised stages:


uncertainty
→working group
→framework
→policy
→rubric
→mandatory online module.


Then, three months later, everyone receives an email announcing that circumstances have changed and a revised framework will shortly be developed to replace the framework that was developed to deal with the previous circumstances.


Trevor spoke. “Perhaps you’re trying to close the question too early.”


Prudence stared at the laptop.


“You keep asking what LLMs are,” Trevor said.


“We need to understand them.”


“Do you?”


“Obviously.”


“Before studying what happens when people work with them?”


Prudence hesitated.


Trevor continued. “You ask whether LLMs think. Whether they understand. Whether they are tools, tutors, cheating devices. Whether they improve learning.”


“Yes.”


“These sound less like research questions than requests for uncertainty to please sit down and behave.”


Charlie smiled.


Trevor continued. “You’ve encountered something rather strange and responded by deciding which existing drawer to put it in.”


“That’s unfair,” Prudence replied.


“Uh huh,” said Trevor. “You also want to label the drawer.”


Charlie drew a box around What happened? on the whiteboard


“Maybe the interesting stuff is small-r research.” he offered.


Barry looked suspicious. “Small-r?”


“Research before Research gets hold of it.”


Barry looked even more suspicious. Trevor continued, “Try something. Notice something. Change something. Try again.”


“That isn’t rigorous.” Barry said with authority.


“Neither is asking 438 teachers whether they strongly agree that ‘AI will play an important role in the future of education’.” Trevor paused. “For the record, 73.6% strongly agree.”


Barry looked impressed.


“I made that up,” said Trevor. 


Barry closed SPSS.


Charlie continued.


Give the bot a job. A specific one. Critic. Translator. Naive reader. Counter-argument generator. Pattern finder. Awkward colleague.”


Trevor offered, “I am particularly strong at awkward colleague.”


Charlie continued, “Then see what happens. What did you delegate? What did you still have to judge? What became easier? What became harder? What did you refuse?”


Prudence was writing now. Charlie found himself talking to Trevor. “What new capacities appeared?”


“Yes.”


“And which disappeared?”


“Yes.”


“And for whom?”


“Exactly.”


Kylie looked thoughtful. “So instead of asking whether teachers have AI competence…”


“…watch what competent action looks like in particular situations.”


“Instead of whether AI improves learning…”


“…ask what happens to the activity when AI joins it.”


“Instead of barriers to adoption…”


“…ask why adoption was assumed to be the destination.”


Barry reopened SPSS, largely because he found the blank screen emotionally difficult.


The meeting was now dangerously close to producing an interesting question and Prudence sensed this.


“We could call it the Situated Generative AI Inquiry Framework.”


“No,” said Charlie.


“The SGAIIF.”


“No.”


“A maturity model?”


“No.”


“A toolkit?”


“No.”


“A microcredential?”


Trevor shut himself down.


The minutes later recorded that the Working Group had enjoyed a “rich and productive discussion.”


This is what university minutes say when nobody has died.


It was agreed that further work was needed to develop a comprehensive framework for responsible, evidence-informed, human-centred, pedagogically appropriate engagement with generative artificial intelligence.


Charlie’s questions did not appear in the minutes. Someone had photographed the whiteboard. That might have been enough. Because perhaps the useful response to LLMs, for now, is not to decide what they are. It is to watch what happens when they are put to work in particular arrangements of people, purposes, rules, knowledge and machines.


Try something.


Notice what moves.


Ask what became possible.


Ask what disappeared.


Ask who had to change.


Then try again.


Small-r research.


A modest proposal, certainly. It has no framework, no maturity model and no six-level rubric. It does, however, have questions. For the moment, that may be the more serious option.



Bibs & bobs #44

  The Question Machine The arrival of large language models created an educational emergency. This was fortunate, because universities are e...