The Chatbot Did Not Break Assessment. It Followed the Instructions
Universities are very good at being surprised by things they have spent years preparing for. They standardise an activity, define the acceptable moves, specify the outputs, publish the criteria, build a workflow around it, put it in a learning management system, and then express alarm when a machine proves rather good at the resulting game. A working group is usually formed at this point, partly to investigate what has happened and partly to ensure that nobody notices what has happened.
Venkatesh Rao has a useful idea for thinking about this. He argues that much of human progress involves making parts of the world more playable. A playable domain has some kind of state, some available moves, some feedback, and some way of telling whether one move worked better than another. Chess is highly playable. A family argument at Christmas is less so.
Mathematics became more playable through notation, formal systems, libraries and proof verification. Programming became more playable through languages, compilers, tests and debuggers.The trick is often not to make the machine smarter, but to simplify and structure the task until even a fairly stupid machine can do it.
This seems rather important for education. Perhaps the interesting question is not: How did LLMs become so good at educational tasks? It is: What did we do to educational tasks to make them so playable?
Quite a lot it seems.
The essay-shaped game board
Consider the university essay. There is a prompt. There is a genre. There is a rubric. There is a word count. There are usually examples. There may be suggested headings. There is often a fairly narrow range of acceptable ways of sounding intelligent.
The student submits some text. The marker compares the text with the rubric. This is less like the wilderness than universities sometimes imagine. It is a game.
Then a LLM arrives. It has “seen” millions of examples of the genre. It knows the moves. It can produce introductions, counterarguments, cautious conclusions and the phrase “further research is needed” without even experiencing the tiny spiritual death normally associated with writing it. Naturally it can play. The surprising thing would be if it could not.
Yet much of the early response to generative AI treated this as though an alien intelligence had broken into assessment through a ventilation shaft. It has not. We have installed the doors and the signage. And, in many cases, a rubric explaining exactly what to do once inside.
Jagged machines, jagged worlds
Some folk, including me, talk about the jaggedness of LLM capability. A LLM can write competent Python and then become confused by a timetable. It can explain relativity and miscount the letters in a word. It can pass an exam and fail at a sandwich. But machine capability is only half the landscape.
Human practices are jaggedly playable too. Coding is highly playable because it has syntax, compilers, tests and error messages. Formal mathematics has verification. Some forms of professional writing have highly stable genres. Other activities are much less obliging.
For instance, knowing when a student is confused but pretending not to be. Or realising that the question being asked is not the problem that needs solving. Or perhaps, working out whether an odd response is wrong, original, funny or all three. Even helping someone formulate a question they do not yet know how to ask.
These practices do not have fixed boards. The state changes while you are looking at it. The criteria move. Sometimes the players change the game, which is one reason statements like “LLMs will replace teachers” are so relentlessly dull. They treat teaching as one task. Teaching is not one task. It is a bag containing dozens of different practices, some of which are highly playable and some of which behave more like ferrets.
We domesticated education first
The usual story says that LLMs arrived and “impacted” education. This gives LLMs rather too much credit. Education had been preparing the ground for decades. We standardised curriculum. We wrote learning outcomes. We aligned assessment. We built rubrics. We created templates. We wrote study guides. We put all of this inside platforms. Then we developed systems for quality assurance, moderation and reporting. None of this was done for LLMs.
It was done for consistency, accountability, scale and the ancient institutional desire to turn difficult things into boxes. But every act of standardisation also makes a practice more legible. And legibility is very useful to machines. Once a practice has clear inputs, conventional outputs and repeatable criteria, a LLM has somewhere to to do its stuff. This is not necessarily a bad thing.
The important point is that machine capability does not appear in isolation. It emerges from an arrangement. A LLM paired with a rubric has capacities it does not have alone. A teacher paired with a LLM has capacities they did not have alone. A student paired with a LLM, an assignment brief, exemplars, a marking guide and three years of institutional formatting has quite a formidable little assemblage.
The useful unit of analysis is not the LLM but student–LLM–teacher–rubric–curriculum–platform–institution. The abilities belong to the arrangement. Bruno Latour would be entirely unsurprised. The speed bump makes the driver slow down. The rubric makes the LLM look educated.
The dangerous assumption
Rao’s idea becomes more interesting when we ask: Should everything become more playable? Education technology has a long-standing fondness for the answer “yes”. If we can specify the learner state, the system can recommend the next move.If we can define the outcome, the system can optimise towards it. If we can collect enough data, the system can personalise the pathway. If we can turn the whole thing into a dashboard, senior management can admire it from a safe distance.
The difficulty is that making something playable always involves deciding what counts. Questions like: What is the state? What is a valid move? What is success? What gets thrown away? Those decisions create the game board. They turn a messy practice into a more manageable representation of that practice. The more stable and explicit that representation becomes, the easier it is for a machine to operate inside it. The machine has not necessarily become cleverer, we have made the world easier for it to read. And, eventually the representation starts to replace the thing represented.
The student becomes a profile. Learning becomes progress against outcomes. Engagement becomes clicks. Writing becomes rubric satisfaction. Understanding becomes successful production. Then, having carefully translated education into machine-readable signals and to our surprise, we discover that machines can read them. This is less a revolution than a filing cabinet receiving a software update.
Keep some missing squares
Perhaps some educational practices should remain only partly playable. Not mystical or confusing, just incompletely specified.
So a good seminar has structure but may not have a fully determined path. A good research problem has constraints but not known moves. A good assessment may define standards without defining every acceptable response. A good teacher sometimes changes the activity because the student has done something unexpected. This is where Christopher Alexander’s pattern language approach may be useful.
A pattern does not say when X happens, do Y. It says: In situations like this, a recurring tension appears. Here is a promising move. Adapt it. That is structure without closure. Playable enough to support action. Not playable enough to remove judgement. Education probably needs more of that.
Maybe assessment should become worse at being a game
Much assessment redesign in the age of LLMs has focused on making tasks harder to game. Add an oral. Add reflection. Add process logs. Add invigilation. Ask for drafts. Require students to explain what they did. Maybe these can help. They all add work for student and teacher.
These responses leave the structure of the game largely intact. They add more moves, more checkpoints and more policing, rather than asking whether we built the wrong game in the first place. The better question is not whether the assessment can still be cheated. It is whether the assessment was too game-like in the first place.
If a LLM can produce an excellent answer, perhaps the first question should not always be: How do we stop the LLM producing it? Maybe it can be: Why did we think producing this answer was good evidence of learning? This is awkwardly less convenient. Which is usually a sign that the question has survived contact with reality.
An assessment may become less trivially playable by asking students to make consequential choices, respond to changing conditions, work with local evidence, justify refusals, revise under challenge, or explain why they rejected plausible machine output. Not because these things are LLM-proof. Nothing involving text is going to be reliably LLM-proof. The point is that they require judgement under conditions that cannot be completely specified beforehand. That is educationally interesting even if the machines all go home tomorrow.
Judgement traces, not surveillance archaeology
Institutions worried about LLMs tend to develop an archaeological interest in student behaviour. We see demands like: show us your drafts, show us your prompts, tell us what model you used, show us when you typed each sentence. Soon the student's cat will be required to provide an independent statement of authenticity.
A judgement trace asks something different. What did you decide? What changed your mind? What did you reject? What did the LLM suggest that you refused? Where were you uncertain? What counted as evidence? That is not an attempt to reconstruct every move.It is evidence that the learner noticed there were moves. A surveillance trace asks: Did you follow the approved path? A judgement trace asks: Can you show where judgement occurred? Those are rather different educational questions.
Refusal matters too
Sometimes the right design move is not to make something more playable. So it's not a good idea to automate this or optimise that. Don’t turn a relationship into a workflow. And don’t convert this uncertainty into a score because someone has discovered a dashboard.
A refusal point is simply a place where making something easier for a machine to handle may make the practice worse. What these practices have in common is that their value depends on interpretation, responsiveness and not knowing in advance exactly what the right move will be. Be it mentoring, feedback, pastoral conversations, or interpretive judgement, for example. Or more importantly, those awkward moments when a learner is still working out what they think and the most helpful response is not yet obvious to anyone involved.
We do not need to claim that humans possess some mysterious quality that silicon can never acquire. That argument usually begins with consciousness and ends somewhere near poetry. The simpler point is that some practices matter because they are ambiguous, responsive and negotiated as they unfold. Remove too much of that mess, and you may indeed make the practice easier for a machine. You may also remove the part that made it worth doing.
AI sensibility may be knowing where the board ends
This argument suggests a more useful account of what is often called AI literacy. It is not prompt tricks, memorising which chatbot currently has the purple button, or developing a working knowledge of menus that will be redesigned before the professional development session has finished. Those things may be useful for a while. So was knowing where the fax paper went.
A more durable form of AI literacy would be knowing how to look at a practice and ask what has made it easy for a LLM to operate there. Ask: What has been standardised? What has been turned into a template? What counts as a successful move? What is easy to verify? What has been left out in order to make the task tidy enough to automate? And, just as importantly, where does the tidiness begin to damage the practice? All of this requires less knowledge about the current chatbot in use and more judgement about the activity itself.
The useful question is not simply, “How do I use this tool?” It is, “What kind of game have we built here, and do we really want the machine to play all of it?”
AI sensibility might mean learning to recognise the playability of a practice. Where are the rules clear? Where are they provisional? Where is verification strong? Where are the criteria contested? What has been compressed? What disappeared in the compression? What capacities does the LLM gain from the surrounding infrastructure? What capacities do humans gain? What capacities do they quietly hand over? And importantly, where should we refuse to finish building the board? That seems more sensible than “using a LLM effectively”. It may even survive the next product launch.
The final institutional surprise
Universities have spent decades making themselves more legible with the support of the ubiquitous LMS. Everything must align. Everything must map. Everything must be measurable. Every aspiration eventually becomes a field in a database. Then LLMs arrive and navigates these structures with suspicious ease. We conclude that the machine has become astonishingly intelligent. Perhaps. But another explanation deserves attention.
We have spent years making large parts of education extremely easy to play. The machines are getting better. So are the game boards. The more useful question is not simply: What can a LLM now do? It’s: What had to happen to this practice before a LLM could become good at it? Or perhaps the more troubling: What do we lose when we make a practice sufficiently playable for a LLM?
Those questions put the attention back on the worlds we have built. Which is unfortunate. It was much more comfortable when the problem was the chatbot.