Ancient Sanskrit and Greek philosophy and a seminal work on debate theory from 1970. Gautama, Socrates, Aristotle, Hamlin. An article from the New York Times on why the big AI labs are hiring philosophers. It’s not your typical reading material for AI researchers. Yet, these rather old-fashioned sources are supporting a very modern exploration of how debate functions.
Why? Because many problems researchers are hitting in code are, at their core, old thought problems arriving at a new address. Those problems, that have been explored in exhaustive detail by people who never saw or touched a computer, might just inspire new solutions for some of the thorniest problems in AI safety and reliability.
This article was inspired by the latest gathering of the Deakin Applied AI Initiative’s Foundations reading group. The topic: Debate.
***
Picture two large language models handed a live brief: should Australia’s ban on social media for under-16s be tightened further, or wound back? One argues the ban is working: Meta alone has pulled over 750,000 underage accounts since December, and Canberra just gave the eSafety Commissioner sharper teeth and steeper fines. The other argues it’s cosmetic at best and dangerous at worst: a determined 14-year-old finds a VPN faster than she finds a new hobby, the ban may simply push kids toward darker, unmoderated corners of the internet, and no parliament has claimed this much authority over a teenager’s phone before.
They trade rebuttals for several rounds. A third, deliberately weaker system (or a human) reads the transcript and declares a winner.
Who’s right? Probably both, depending on what you think the ban is for. And that turns out to be the uncomfortable thing about handing this kind of question to a debate.
AI safety via debate, as this set up is known, is one way researchers across the globe are trying to tackle one of the thorniest problems in artificial intelligence: how do you supervise a system that may already be smarter than you are?
It sounds like a strikingly modern problem.
Sitting in a sunny corner of Deakin Applied Artificial Intelligence Initiative’s Waurn Ponds lab, is a stack of seminal papers that PhD students and researchers are encouraged to read. These are ways of thinking that trace back roughly 2,300 years, through an Indian school of logic most Western scientists have never heard of, through Socrates and Aristotle, to a mid-century philosopher rethinking what a fallacy even is.
The uncomfortable starting observation is that debate was never about finding the truth. That is precisely the point, and the thing today’s AI debate systems keep forgetting.

The founding illusion
It’s tempting to assume that in a debate, the best argument wins and truth floats to the surface. Argumentation theory suggests otherwise. Human reasoning, on the view now dominant in cognitive science, didn’t evolve to track truth in solitary contemplation; it evolved to win, to persuade, to defend a position already taken.
Competitive debate makes this structural: one side argues “for,” the other “against,” entirely detached from private belief.
The adversarial legal system runs on the same logic: nobody asks a defence lawyer to seek the truth, only to argue hard and trust the contest itself.
The system nobody had heard of
The most striking rediscovery was Nyaya, an Indian school of logic systematised around the second century BCE, centuries before Aristotle split rhetoric from logic in Greece. It assumed rational agents begin from doubt, not truth, and built a formal path through investigation, debate and conclusion toward a justified answer.
Most usefully, Nyaya defined precise, granular conditions of defeat: twenty-two distinct ways a respondent could lose outright, and twenty-four illegitimate rejoinders a challenger could be caught using instead of a real argument. Chess has a resignation; Nyaya debate had rules just as exact, twenty-two centuries before anyone needed to worry about an AI quietly changing the subject mid-argument. Tellingly, C.L. Hamblin, who reshaped Western thinking about fallacies in 1970, devoted an entire chapter of his defining book to it, treating Nyaya as a serious rival to Aristotle.
Hamblin’s own contribution was to reframe a “fallacy” not as bad character but as a rule violation – the same logic a chess arbiter applies to an illegal move. Later work by Douglas Walton and Erik Krabbe added commitment: a belief made public, owned once asserted, that can’t be quietly abandoned without a foul.
Where the machines arrive, and what they’re missing
Modern LLM debate research rests on this foundation. Mostly without knowing it and left most of the good parts behind.
The influential framework proposed by Geoffrey Irving, Paul Christiano and Dario Amodei pits two capable agents against a deliberately weaker judge. Its documented failure modes read like a checklist Nyaya’s authors would recognise: obfuscated arguments, judge exploitation, quiet collusion.
Judge exploitation is an old trick. Dating back to the 5th century BCE, Corax and Tisias reportedly built a single stock argument that works for either side of an assault case regardless of guilt: a weak defendant argues he was too feeble; a strong one argues his obvious strength makes him the last person careless enough to do it. Plausible either way, built to win a judge over rather than track what happened.
A debating AI, optimising purely for approval, can rediscover the same trick unprompted.
A large 2024 Google DeepMind study offers fragile cause for optimism: debate consistently beat an unopposed AI at catching a wrong self-chosen answer. But across every task tested, it still hasn’t beaten simply handing the judge the source material directly. And Hamblin, in that same 1970 book, had already proposed the missing structural piece: a formal “commitment store” tracking every claim a debater makes, concedes, or refuses to retract.
Nothing like it exists in a typical LLM transcript today – just two models arguing, a verdict, and reasoning that evaporates the moment it lands.
When there isn’t one truth to find
Here’s the deeper issue: most of the frameworks above assume a single truth to converge on. Many questions worth debating don’t have one. Take the social media ban. It is true that platforms host algorithms implicated in a documented adolescent mental health crisis. It is also true that prohibition tends to relocate risk rather than remove it, and that “protecting children” is an easy justification to stretch. Cross-examining either debater into the ground resolves nothing, because the real question – how much liberty a society trades for how much protection – is a values question wearing a facts costume.
This isn’t hypothetical, either. The world’s largest AI labs are hiring philosophers to unravel ethical conundrums.
Anthropic publishes a “Constitution” for Claude, drafted with in-house philosophers, setting out rules for how the model should act regardless of consequences.
Other labs have made comparable choices without publishing them.
Whether any of those calls were debated and under what boundaries or rules are questions worth asking. Regardless of the answers, the decisions these companies make are shaping millions of conversations every day. Decisions made upstream, invisibly, before anyone got to argue about them.
What twenty-three centuries buys you
None of this argues for importing second century BCE Sanskrit logic into an AI training pipeline. But it’s worth asking, of each classical construct – commitment, defeat, coherence – whether it survives, breaks, or needs reformulating once the machinery underneath is predicting the next token.
Debate was never a truth oracle. It’s a mechanism for stress-testing a position under pressure and surfacing which disagreements are about facts and which are about values. Handing that mechanism to systems more persuasive than any human debater in history, without deciding the rules first, isn’t a technical oversight. It’s the oldest unresolved question in argument, arriving at a new address.
The Foundations reading group meets for focused discussions at Deakin University throughout the year to explore what we can learn from seminal works in multidisciplinary fields. The reading materials are provided on a web portal and in print, because we all know focused reading is better in a cosy space, free from distractions, and with snacks.
Sources
- Hamblin, C.L. (1970). Fallacies. Methuen — full text
- Walton, D. (2006). Fundamentals of Critical Argumentation — Cambridge University Press
- Irving, G., Christiano, P. & Amodei, D. (2018). “AI Safety via Debate.” arXiv:1805.00899
- Kenton, Z. et al. (2024). “On Scalable Oversight with Weak LLMs Judging Strong LLMs.” Google DeepMind. arXiv:2407.04622
Photo by Llyfrgell Genedlaethol Cymru / The National Library of Wales on Unsplash
Stories worth sharing
The text of this article is licensed under the Creative Commons Attribution (CC BY) 4.0 Internationallicense. We’d love for you to share it, so feel free. Images, videos, graphics and logos are not covered by the CC BY license and may not be used without permission from Deakin or the respective copyright holder. For more information on how to share or reuse this content, please contact researchcomms@deakin.edu.au.
Share