September 06, 2026

Bibs & bobs #47

The Chatbot Did Not Break Assessment. It Followed the Instructions


Universities are very good at being surprised by things they have spent years preparing for. They standardise an activity, define the acceptable moves, specify the outputs, publish the criteria, build a workflow around it, put it in a learning management system, and then express alarm when a machine proves rather good at the resulting game. A working group is usually formed at this point, partly to investigate what has happened and partly to ensure that nobody notices what has happened.


Venkatesh Rao has a useful idea for thinking about this. He argues that much of human progress involves making parts of the world more playable. A playable domain has some kind of state, some available moves, some feedback, and some way of telling whether one move worked better than another. Chess is highly playable. A family argument at Christmas is less so.


Mathematics became more playable through notation, formal systems, libraries and proof verification. Programming became more playable through languages, compilers, tests and debuggers.The trick is often not to make the machine smarter, but to simplify and structure the task until even a fairly stupid machine can do it. 


This seems rather important for education. Perhaps the interesting question is not: How did LLMs become so good at educational tasks? It is: What did we do to educational tasks to make them so playable?


Quite a lot it seems.


The essay-shaped game board


Consider the university essay. There is a prompt. There is a genre. There is a rubric. There is a word count. There are usually examples. There may be suggested headings. There is often a fairly narrow range of acceptable ways of sounding intelligent.


The student submits some text. The marker compares the text with the rubric. This is less like the wilderness than universities sometimes imagine. It is a game.


Then a LLM arrives. It has “seen” millions of examples of the genre. It knows the moves. It can produce introductions, counterarguments, cautious conclusions and the phrase “further research is needed” without even experiencing the tiny spiritual death normally associated with writing it. Naturally it can play. The surprising thing would be if it could not.


Yet much of the early response to generative AI treated this as though an alien intelligence had broken into assessment through a ventilation shaft. It has not. We have installed the doors and the signage. And, in many cases, a rubric explaining exactly what to do once inside.


Jagged machines, jagged worlds


Some folk, including me, talk about the jaggedness of LLM capability. A LLM can write competent Python and then become confused by a timetable. It can explain relativity and miscount the letters in a word. It can pass an exam and fail at a sandwich. But machine capability is only half the landscape.


Human practices are jaggedly playable too. Coding is highly playable because it has syntax, compilers, tests and error messages. Formal mathematics has verification. Some forms of professional writing have highly stable genres. Other activities are much less obliging.


For instance, knowing when a student is confused but pretending not to be. Or realising that the question being asked is not the problem that needs solving. Or perhaps, working out whether an odd response is wrong, original, funny or all three. Even helping someone formulate a question they do not yet know how to ask.


These practices do not have fixed boards. The state changes while you are looking at it. The criteria move. Sometimes the players change the game, which is one reason statements like “LLMs will replace teachers” are so relentlessly dull. They treat teaching as one task. Teaching is not one task. It is a bag containing dozens of different practices, some of which are highly playable and some of which behave more like ferrets.


We domesticated education first


The usual story says that LLMs arrived and “impacted” education. This gives LLMs rather too much credit. Education had been preparing the ground for decades. We standardised curriculum. We wrote learning outcomes. We aligned assessment. We built rubrics. We created templates. We wrote study guides. We put all of this inside platforms. Then we developed systems for quality assurance, moderation and reporting. None of this was done for LLMs.


It was done for consistency, accountability, scale and the ancient institutional desire to turn difficult things into boxes. But every act of standardisation also makes a practice more legible. And legibility is very useful to machines. Once a practice has clear inputs, conventional outputs and repeatable criteria, a LLM has somewhere to to do its stuff. This is not necessarily a bad thing.


The important point is that machine capability does not appear in isolation. It emerges from an arrangement. A LLM paired with a rubric has capacities it does not have alone. A teacher paired with a LLM has capacities they did not have alone. A student paired with a LLM, an assignment brief, exemplars, a marking guide and three years of institutional formatting has quite a formidable little assemblage.


The useful unit of analysis is not the LLM but student–LLM–teacher–rubric–curriculum–platform–institution. The abilities belong to the arrangement. Bruno Latour would be entirely unsurprised. The speed bump makes the driver slow down. The rubric makes the LLM look educated.


The dangerous assumption


Rao’s idea becomes more interesting when we ask: Should everything become more playable? Education technology has a long-standing fondness for the answer “yes”. If we can specify the learner state, the system can recommend the next move.If we can define the outcome, the system can optimise towards it. If we can collect enough data, the system can personalise the pathway. If we can turn the whole thing into a dashboard, senior management can admire it from a safe distance.


The difficulty is that making something playable always involves deciding what counts. Questions like: What is the state? What is a valid move? What is success? What gets thrown away? Those decisions create the game board. They turn a messy practice into a more manageable representation of that practice. The more stable and explicit that representation becomes, the easier it is for a machine to operate inside it. The machine has not necessarily become cleverer, we have made the world easier for it to read. And, eventually the representation starts to replace the thing represented.


The student becomes a profile. Learning becomes progress against outcomes. Engagement becomes clicks. Writing becomes rubric satisfaction. Understanding becomes successful production. Then, having carefully translated education into machine-readable signals and to our surprise, we discover that machines can read them. This is less a revolution than a filing cabinet receiving a software update.


Keep some missing squares


Perhaps some educational practices should remain only partly playable. Not mystical or confusing, just incompletely specified.


So a good seminar has structure but may not have a fully determined path. A good research problem has constraints but not known moves. A good assessment may define standards without defining every acceptable response. A good teacher sometimes changes the activity because the student has done something unexpected. This is where Christopher Alexander’s pattern language approach may be useful.


A pattern does not say when X happens, do Y. It says: In situations like this, a recurring tension appears. Here is a promising move. Adapt it. That is structure without closure. Playable enough to support action. Not playable enough to remove judgement. Education probably needs more of that.


Maybe assessment should become worse at being a game


Much assessment redesign in the age of LLMs has focused on making tasks harder to game. Add an oral. Add reflection. Add process logs. Add invigilation. Ask for drafts. Require students to explain what they did. Maybe these can help. They all add work for student and teacher.


These responses leave the structure of the game largely intact. They add more moves, more checkpoints and more policing, rather than asking whether we built the wrong game in the first place. The better question is not whether the assessment can still be cheated. It is whether the assessment was too game-like in the first place.


If a LLM can produce an excellent answer, perhaps the first question should not always be: How do we stop the LLM producing it? Maybe it can be: Why did we think producing this answer was good evidence of learning? This is awkwardly less convenient. Which is usually a sign that the question has survived contact with reality.


An assessment may become less trivially playable by asking students to make consequential choices, respond to changing conditions, work with local evidence, justify refusals, revise under challenge, or explain why they rejected plausible machine output. Not because these things are LLM-proof. Nothing involving text is going to be reliably LLM-proof. The point is that they require judgement under conditions that cannot be completely specified beforehand. That is educationally interesting even if the machines all go home tomorrow.


Judgement traces, not surveillance archaeology


Institutions worried about LLMs tend to develop an archaeological interest in student behaviour. We see demands like: show us your drafts, show us your prompts, tell us what model you used, show us when you typed each sentence. Soon the student's cat will be required to provide an independent statement of authenticity.


A judgement trace asks something different. What did you decide? What changed your mind? What did you reject? What did the LLM suggest that you refused? Where were you uncertain? What counted as evidence? That is not an attempt to reconstruct every move.It is evidence that the learner noticed there were moves. A surveillance trace asks: Did you follow the approved path? A judgement trace asks: Can you show where judgement occurred? Those are rather different educational questions.


Refusal matters too


Sometimes the right design move is not to make something more playable. So it's not a good idea to automate this or optimise that. Don’t turn a relationship into a workflow. And don’t convert this uncertainty into a score because someone has discovered a dashboard.


A refusal point is simply a place where making something easier for a machine to handle may make the practice worse. What these practices have in common is that their value depends on interpretation, responsiveness and not knowing in advance exactly what the right move will be. Be it mentoring, feedback, pastoral conversations, or interpretive judgement, for example. Or more importantly, those awkward moments when a learner is still working out what they think and the most helpful response is not yet obvious to anyone involved.


We do not need to claim that humans possess some mysterious quality that silicon can never acquire. That argument usually begins with consciousness and ends somewhere near poetry. The simpler point is that some practices matter because they are ambiguous, responsive and negotiated as they unfold. Remove too much of that mess, and you may indeed make the practice easier for a machine. You may also remove the part that made it worth doing.


AI sensibility may be knowing where the board ends


This argument suggests a more useful account of what is often called AI literacy. It is not prompt tricks, memorising which chatbot currently has the purple button, or developing a working knowledge of menus that will be redesigned before the professional development session has finished. Those things may be useful for a while. So was knowing where the fax paper went.


A more durable form of AI literacy would be knowing how to look at a practice and ask what has made it easy for a LLM to operate there. Ask: What has been standardised? What has been turned into a template? What counts as a successful move? What is easy to verify? What has been left out in order to make the task tidy enough to automate? And, just as importantly, where does the tidiness begin to damage the practice? All of this requires less knowledge about the current chatbot in use and more judgement about the activity itself.


The useful question is not simply, “How do I use this tool?” It is, “What kind of game have we built here, and do we really want the machine to play all of it?”


AI sensibility might mean learning to recognise the playability of a practice. Where are the rules clear? Where are they provisional? Where is verification strong? Where are the criteria contested? What has been compressed? What disappeared in the compression? What capacities does the LLM gain from the surrounding infrastructure? What capacities do humans gain? What capacities do they quietly hand over? And importantly, where should we refuse to finish building the board? That seems more sensible than “using a LLM effectively”. It may even survive the next product launch.


The final institutional surprise


Universities have spent decades making themselves more legible with the support of the ubiquitous LMS. Everything must align. Everything must map. Everything must be measurable. Every aspiration eventually becomes a field in a database. Then LLMs arrive and navigates these structures with suspicious ease. We conclude that the machine has become astonishingly intelligent. Perhaps. But another explanation deserves attention.


We have spent years making large parts of education extremely easy to play. The machines are getting better. So are the game boards. The more useful question is not simply: What can a LLM now do? It’s: What had to happen to this practice before a LLM could become good at it? Or perhaps the more troubling:  What do we lose when we make a practice sufficiently playable for a LLM?


Those questions put the attention back on the worlds we have built. Which is unfortunate. It was much more comfortable when the problem was the chatbot.



 

August 24, 2026

Bibs & bobs #46

 We Designed the Habitat

I have been reading Owen Jones’s Force of Nature [1], a book about natural selection which has the unfortunate side effect of making almost everything look like natural selection. This is probably why evolutionary biologists should not be allowed near university strategic planning.


Jones’s argument is not the tired one that evolution happened a very long time ago and eventually produced us, universities and the committee meeting. His point is that selection is happening all the time. More importantly, what we do changes the selection pressures acting in the world. This turns out to be an awkward thought to have while watching education respond to LLMs.


The Sin of Assumed Invariance


Jones uses commercial fishing to make the problem clear. Fishing industries preferentially catch large fish. Yet they have often behaved as though fish populations will remain more or less the same: remove this year’s large fish and nature will kindly manufacture another standard batch for next year.


Jones calls this the Sin of Assumed Invariance. The problem is that the fishing is itself changing the population. If being large makes you more likely to end up on a plate, being smaller becomes rather a good reproductive strategy. Jones describes evidence that intense size-selective fishing can alter growth, maturation and reproduction surprisingly quickly. The intervention changes the thing being intervened upon.


Education appears to have adopted the Sin of Assumed Invariance as an implementation strategy. We introduce AI detectors, declarations of use, supervised assessment, oral defences, locked browsers and diagrams explaining precisely when ChatGPT may be consulted. Then we ask whether the intervention worked. Somewhere inside this model sits a strangely motionless student. The policy changes. The assessment changes. The LLM changes. The student, apparently, waits. Unfortunately, students learn. So do teachers.


Introduce an AI detector and you have not simply detected AI use. You have made undetectable AI use more valuable. Introduce an AI declaration and you have not simply produced transparency. You have created a new genre: the acceptable account of how AI was used. Invent an “AI-proof assessment” and there will shortly be students discovering how a LLM can help them complete it.


This is not evidence that students have suddenly become morally defective. It is evidence that environments have consequences. Jones gives us a much better question than: Did our intervention work? He suggests that we ask instead: What did our intervention select for?


That question should probably be printed above the door of every university AI working group. Unfortunately, there is usually already something above the door. Future Ready, perhaps. Or Responsible AI. Possibly Transforming Learning for an AI-Enabled Future, if the sign was commissioned by consultants and the available wall space was generous.


These phrases perform an important institutional function. They imply that somewhere there is a future, that it has already been inspected, and that readiness consists largely of arriving there in the correct attire.


Jones suggests a less reassuring possibility. Every intervention alters the environment in which subsequent behaviour occurs. The policy does not merely regulate practice; it becomes part of the conditions to which practice adapts. The assessment rule, detector, declaration form and approved-use matrix all enter the habitat and begin exerting selection pressures of their own.


So the awkward question for the working group is not simply whether its policy encourages “responsible AI use.” It is whether the policy makes responsible use more viable than irresponsible use, or merely makes responsible-looking use more viable. Responsible-looking use sharpens the Jones point: institutional interventions can select for the appearance of the desired behaviour rather than the behaviour itself.


Future Ready is much neater. It also fits on a lanyard.


We built the fitness function


Jones becomes even more useful when he turns to evolutionary computation. The basic trick is wonderfully simple. Generate lots of possible solutions. Introduce variation. Test them. Let better-performing solutions contribute to the next generation. Repeat.


But there is a rather important detail. Someone has to specify what counts as “better.” Evolutionary computation therefore requires what Jones calls a fitness function: a specification of the problem and the characteristics against which candidate solutions will be judged.


Education already has these. We call them rubrics. Suppose an assignment rewards coherent prose, clear structure, plausible argument, appropriate referencing and something identifiable as critical thinking. Then a technology arrives that can produce coherent prose, clear structure, plausible argument and references while sounding sufficiently thoughtful to survive moderate exposure to a marking rubric.


Education responds with astonishment. This is rather like constructing a bird feeder and expressing outrage when birds turn up. We designed the habitat. Then we complained about the wildlife. The problem is not necessarily that the fitness function is bad. It is that LLMs expose a distinction we had previously been able to ignore. The declared fitness function might be: develop historical understanding. The operational fitness function might be: produce 2,000 words containing the characteristics that cause a marker to award 73%.


Before LLMs, these could be treated as sufficiently close cousins. Producing the essay generally required enough reading, thinking and writing that the artefact provided some evidence of what had happened inside the student. Not perfect evidence. Education has always survived on proxies. Otherwise assessment would require opening students and inspecting the learning directly, which would create paperwork and sometimes a lot of blood.


LLMs disturb the proxy. They produce some of the valued characteristics of the artefact without necessarily producing the educational process we thought the artefact represented. The machine has not destroyed assessment. It has done something considerably ruder. It has revealed the fitness function.


Then it gets jagged [2]


At this point the evolutionary story can become far too neat. The environment changes. People adapt. Everyone acquires AI literacy. There is a webinar. Problem solved.


But real people are inconveniently jagged. A student can be superb at getting useful responses from a LLM and terrible at recognising nonsense in those responses. Another can possess excellent disciplinary judgement but use ChatGPT as though it were Google with more adjectives. A teacher can understand their students and subject extraordinarily well while having almost no idea what current LLMs can do. Another can prompt magnificently while remaining slightly uncertain what any of it is for. Calling all this “AI literacy” gives the comforting impression of a ladder. Jaggedness gives us something more like a mountain range designed by a committee that disagreed about gravity.


People occupy different adaptive positions along multiple dimensions at once. That means there is no single route from novice to expert. More importantly, moving is difficult.


The trouble with leaving somewhere that works


Jones describes another problem in evolutionary computation. Candidate solutions can become too similar and converge on a local sub-optimal solution. His less impressive but much more useful translation is: stuck in a rut. The phrase is unfair to educational practice because ruts are often extremely efficient.


Consider the essay. The teacher knows how to set it. Students know roughly what one is. The LMS knows where to put it. The rubric knows how to judge it. The moderation process recognises it. The accreditation documents contain reassuring boxes into which it fits. Nobody has to explain to Quality Assurance why students are constructing imaginary conversations between Napoleon and a chatbot.


This arrangement may not represent the finest educational possibility available to humanity. But it works. And that matters. When people are told that LLMs require them to move to a new adaptive position, we frequently forget that the journey may first make them less competent.


A teacher who has refined an assessment for ten years replaces it with something experimental. Workload rises. Predictability falls. Students become confused. Moderators become interested. We can assume that nobody wants moderators to become interested.


So “resistance to change” may sometimes be resistance. But sometimes it is perfectly sensible behaviour by someone well adapted to their current environment who can see that getting somewhere supposedly better requires crossing a stretch of territory in which things get worse.


Jaggedness makes that crossing different for everyone. For one teacher, having a LLM challenge students’ explanations is an obvious next experiment. For another, opening ChatGPT is the experiment. For one student, a LLM expands what can be questioned. For another, fluent machine prose closes questioning down because it looks so thoroughly like an answer.


There isn't one transition to AI-enabled education. There are thousands of little transitions, made from different starting points, with different risks, capacities and adjacent possibilities. The arrow in the implementation diagram is therefore lying. It knows this. It simply has excellent graphic design.


Variation may be the thing we need


Jones’s discussion of evolutionary computation contains one final but useful twist. When candidate solutions become too similar, evolutionary algorithms can deliberately increase variation. Diversity helps the system escape the local rut and explore more of the possible solution space. This seems almost exactly opposite to our institutional instinct.


Faced with LLMs, education wants convergence and certainty. Approved practice. Approved tools. Approved prompts. Approved assessment designs. Frameworks explaining how to comply with other frameworks.


Some standardisation is clearly necessary. There are real questions about privacy, equity, intellectual responsibility and assessment validity. But if we do not yet know what good educational practice with LLMs looks like, and I don't think we do, then prematurely selecting one approved destination may be precisely the wrong thing to do.


We need variation. We need to encourage small experiments. We need to applaud cheap failures. We need different teachers trying different things from different starting positions. We should reward students noticing what helps them think and what merely helps them produce. Not because whatever emerges will necessarily be wonderful. Evolution has produced both the human brain and the blobfish. Selection is not a synonym for progress.


Which brings us back to the most important Jones question which is not How do we get education to adapt to AI? But is What is our educational environment selecting for?


If our assessments reward the production of answers, LLMs will flourish there. If our AI policies reward concealment, concealment will improve. If our institutions reward standardisation, safe practices will outcompete interesting ones. And if moving to new practices imposes all the risk on individual teachers while the institution retains all the benefits, staying in the rut may remain an exceptionally fit behaviour.


LLMs may therefore be exposing something more interesting than the weaknesses of students or the limitations of assessment. They may be exposing the selection pressures of education itself.


We built and maintain the habitat. We specified the fitness functions. We rewarded some behaviours and made others expensive. Then a peculiar new creature appeared that was exceptionally well suited to parts of the environment we had constructed.


Perhaps the interesting task is not to domesticate it. Perhaps it is to ask why the habitat looks like this in the first place. Sadly, that question does not fit neatly into the implementation framework. But be reassured, a working group will no doubt be established.


Notes


[1] Jones, O. D. (2026). Force of nature : understanding evolution's deepest logic--and putting it to use (First edition.). W.W. Norton & Company.  


[2] I have written a number of posts about the jaggedness of AI savviness in teachers and students and the problems that flow from that. This is one of them.

August 21, 2026

Bibs & bobs #45

 Fourteen Seconds Before the Worksheet

Large Language Models arrived in education. This was potentially significant. Here was a peculiar new object capable of producing language, arguing, role-playing, translating, explaining, inventing examples, changing perspectives, making unexpected connections, confidently fabricating things and, when challenged, apologising in a manner suggesting absolutely no emotional consequences whatsoever.

Education regarded it carefully. For about fourteen seconds. Then somebody said:


“Can it make a worksheet?”


And that was more or less that.


The domestication programme [1]


Education has considerable experience with new technologies. Whenever something strange arrives, there is a short period during which it might conceivably change what we do. This is dangerous. Fortunately, institutions have developed powerful antibodies. The unfamiliar object is surrounded by committees, frameworks, learning outcomes, risk registers and people asking whether it integrates with the LMS. 


Eventually it is rendered harmless by being made to perform an existing activity slightly faster. The computer became a typing machine. The internet became somewhere to put readings. The learning management system became a filing cabinet that sends email. And the Large Language Model, an object apparently assembled from a substantial portion of recorded human language, became a machine for making lesson plans.


Domestication complete. Nobody was hurt. 


This is an astonishing achievement. It should not be underestimated. We have taken something genuinely odd, a machine with which one can conduct a rapid, recursive conversation about almost anything, and taught it to produce: lesson plans, quizzes, PowerPoint slides, rubrics, model answers, avatars that “teach”, differentiated worksheets, feedback comments, and cheerful little exit tickets asking students what they learned today.


This is rather like discovering a visiting extraterrestrial civilisation and immediately asking whether its spacecraft can laminate. The answer may well be yes. That doesn't make it the interesting question. Yet there is something wonderfully reassuring about the whole process.


If you type: Create a Year 9 lesson for photosynthesis.


Seconds later there will appear a lesson plan, a worksheet, an infographic, three learning outcomes, several activities, six multiple-choice questions and, if the machine has not been adequately supervised, a rubric. This looks enormously productive. There is certainly a lot of it. Quantity has always enjoyed a somewhat undeserved reputation in education.


We have been here before


There is a small historical joke buried in all this. In the early days of educational computing, some teachers wrote software to teach particular concepts. This was no simple matter.


Suppose you wanted to write a program to teach fractions. Before the computer would do anything useful, someone had to decide what a fraction was, what students commonly misunderstood about fractions, which examples mattered, which sequence might help, what counted as an error and what should happen next. This involved an irritating amount of thinking and, an awkward thing sometimes happened.


The person who learned most from the educational software was the person who wrote it. Writing the program forced its author to take the concept apart, examine it, find its difficult edges and put it back together. The student later encountered the finished software and clicked NEXT.


This was not quite the intended distribution of learning. Still, there was something interesting going on. The difficult work of making the teaching material was itself intellectually productive. The problem was solved. The machine can do it.


Cognitive offloading, now with clip art


Ask a LLM: Give me five analogies for entropy suitable for a ten-year-old.


The machine disappears briefly into whatever passes for thought in these circumstances and returns with five analogies. Three are faintly dreadful. One involves a bedroom. There is nearly always a bedroom. One is actually quite good. The good one is copied into PowerPoint. Yay! Success!


But something curious has happened. Someone, or something, has done the wandering. Possible representations have been generated. Comparisons have been made. Explanations have been bent, tested and discarded. The edges of the concept have been bumped into. The human receives analogy number four.


This is called cognitive offloading, which makes the transfer of intellectual activity sound reassuringly like placing luggage in an overhead compartment. Clearly, effort has been saved. Nobody had to spend twenty minutes producing five not-quite-right analogies.


This is generally regarded as a benefit because educational systems have spent decades establishing the important principle that teachers have far too much to do. But there is a small possibility that the five not-quite-right analogies were where some of the learning was hiding. We may have automated the troublesome part that required us to think and retained the PowerPoint slide. This would be mildly funny if it were not such a familiar move.


The photocopier acquires opinions


The odd thing about LLMs is not that they make educational materials. Humans have been making educational materials for centuries. Some civilisations are believed to have collapsed beneath the weight of them. The odd thing is the interaction.


You can say: “No, that's not what I mean” or, “Try another explanation,” or, “Assume the opposite,” or, “Give me an example where this breaks down,” or, “Argue against yourself,” or, “Make the case from a perspective I dislike,” or, “That's nonsense.” And the machine will always comply. It does not sigh. It does not glance at the clock. It does not say that this was covered in last week's meeting. This is a genuinely peculiar new capacity. So naturally, we have given it a template for a Year 8 worksheet.


Perhaps the output isn't the interesting object


One consequence of domestication is that attention settles on the thing produced. The lesson plan, the infographic, the worksheet or the model answer. The finished thing looks educational, so we assume that is where the educational value must be.


But perhaps the interesting object is not the output. Perhaps it is the traffic: the questions, the rejected answers, the sudden change of direction, the moment someone notices that a plausible explanation is wrong, the comparison between three different representations, the argument about what counts as a good example or, the realisation that the question itself is badly worded.


When the machine says something unexpected. The human thinks, effectively, “Hang on.” That begins to look less like content production and more like intellectual activity. Which is inconvenient, because intellectual activity is much harder to upload to the LMS.


The slightly embarrassing possibility


Students already have access to versions of these machines, often better versions and some students may be more LLM skilled than the teacher [2]. This creates a curious arrangement.


The institution may use a LLM privately to produce teaching materials. The teacher may use a LLM privately to produce teaching materials. The student may use a LLM privately to produce the assignment. Everybody then meets in the classroom and pretends the interesting thing is the document. This is a magnificent piece of theatre. The machines converse with everyone backstage while the humans exchange PDFs out front.


Perhaps there is another possibility. Rather than always hiding the interaction and presenting its polished remains, we could sometimes make the interaction itself available for inspection. Not because everyone should learn a set of approved prompting tricks. That would simply be domestication with an advanced settings menu. But because there may be something worth seeing in how different people work with a machine that answers too easily.


You can ask things like: What got accepted? What got challenged? What was ignored? (Ouch) Who noticed the hidden assumption? Who decided that using the machine at all is making matters worse? These are not really LLM skills. They are ways of working with claims, explanations, possibilities and uncertainty. They merely become unusually visible when another participant in the conversation can produce a confident answer in two seconds.


A small problem with responsible use


Educational organisations are understandably concerned that students should use AI responsibly. The usual response is to write a policy. This is rather like responding to the invention of the bicycle by issuing a document on responsible balance. The policy may be necessary. It is unlikely to teach balance.


If working with systems like these becomes part of ordinary intellectual life, then students probably need encounters with actual practices rather than simply instructions about permitted outcomes. They need to see people getting it wrong. They need to see seductive answers rejected. They need to see uncertainty survive contact with fluent prose. They need to see that sometimes the best use of a LLM is to continue the conversation. And sometimes they need to see that the best use is to close the window and go for a walk.


And to think the unthinkable, there may even be no single correct way of working with the machine. This will be disappointing for anyone currently preparing a framework. Ooops!


A successful domestication


The deepest domestication of LLMs may therefore have little to do with banning them or permitting them. It happens when we decide, almost without noticing, what kind of thing they are. Maybe they are seen as a resource generator or productivity aid or a tutor, or cheating device. Maybe a feedback machine.


Once the label sticks, the possibilities begin to shrink. The strange object becomes familiar. The familiar object becomes manageable. The manageable object gets a procurement category. And eventually somebody produces a two-page guide called Five Ways to Use Generative AI Effectively in Your Classroom.


At which point the alien spacecraft has successfully been fitted with a laminator.


But perhaps we should leave the machine strange for a little longer. Not because it is magical. It isn't. Not because it will transform education. Things have been promising to transform education for quite some time and education remains impressively difficult to transform. But because there is something almost heroic about encountering a genuinely unfamiliar technology and resisting, for slightly longer than fourteen seconds, the urge to make it do what we were already doing.


The worksheet can wait.


It has waited before.


Notes


[1] A less playful account can be found in this short paper: Bigum, C. (2023, June). Teacher librarians, orthodox and heterodox: making sense of and playing in a world increasingly run by machines. Access, 37(2), 31-35. https://drive.google.com/file/d/1XJSD294lvIjAAwc6rXkwXK8iA8zx-NJX/view?usp=sharing 


[2] I have a post that points to the problem of the uneven or jagged nature of LLM savviness among students and their teachers. 


Bibs & bobs #47

The Chatbot Did Not Break Assessment. It Followed the Instructions Universities are very good at being surprised by things they have spent y...