Judge the argument, not the room.
AgoraCrux listens to a debate — live, or a recording — and separates what was actually established from what merely performed well. Every score traces back to the words that earned it, and every gap between the two gets named instead of averaged away.
- Evidence-linked
- Every score can produce the utterances behind it — or it renders as “insufficient evidence,” never a number with nothing under it.
- Counted, not guessed
- The model classifies discrete events; arithmetic aggregates them. It's never asked to invent a score.
- Absent isn't zero
- A speaker never asked a question shows “—”. An index not yet built shows “·”. Neither shows 0%.
An analyzer that only confirms what you already believed isn't an analyzer.
It's a rationalization engine, and people notice eventually. The credibility of everything on this page rests on one rule.
The scorer must be capable of returning a result the user didn't expect.
- Speaker 2 dominated the exchange but established less than you'd think.
- Speaker 1's position was not refuted; it was abandoned.
- This format does not support a comparative verdict.
Online debate is corrupted by flawed metrics.
Judges — human or crowd — reward popularity, charisma, interruption, and speaking volume rather than logical rigor and evidence quality. It happens everywhere: academic debate, workplace argument, legal argument, every comment section on the internet.
Chaos is the default
One founder estimated roughly 90% of their debate experiences were chaotic rather than structured — people talking over each other, endless scroll, noise drowning substance. Chaos isn't incidental to bad metrics; it produces them, because whoever is loudest and fastest wins by default.
There's no referee
Moderation is absent, inconsistent, or biased, so disputes have no legitimate resolution. People expect a neutral arbiter for every other competitive interaction in society. Online discourse doesn't have one.
The result: the person who argued better and the person who appeared to win are routinely different people, and afterward nobody can prove it either way.
Three ideas carry this. Everything else is table stakes.
The rest of this page, and the full build status, are proof of engineering. These are why it's built at all.
Two-track scoring
PartialEvery score splits into two tracks, computed completely separately: propositional support — whose position was actually established — and dialectical performance — who discharged their obligations in the exchange. These come apart constantly. A speaker can dominate the room and establish nothing. A speaker can hold the better-grounded case and still have the worse night. Naming that gap is the thing nobody else offers.
Headlines are generated from thresholds on the computed indices, never written freehand — a summarizer would just drift toward whatever the loudest voice asserted. Right now two of the four indices behind it are built, so live output can only speak to responsiveness spread; the rest of the picture is on the way.
Illustrative — the scatter view is designed, not shipped yet.
The confidence veil
Spec’dBefore comparing speakers at all, the system checks whether this was even a comparable exchange: does one party have post-hoc narration the other can't answer? Is the cut non-chronological? Are participants in comparable roles? Is speaking time wildly lopsided? When something fires, a fine diagonal hatch is drawn across the score panel — scores stay legible underneath, but visibly qualified — with a caption naming exactly why.
Nothing else in the interface uses hatching, so its presence alone is information. Three states, not two: no hatch and “all clear” on a clean format, no hatch and “not applicable” where a check simply can't run on this format, or the hatch with the flags named. A clean live debate should never have to look identical to a clean upload it wasn't.
/// SCORES DISCOUNTED 40% /// ambush format · unopposed narration
Concept illustration — designed, not shipped yet.
What it refuses to measure
Live — as policyComposure, tone, and affect are not scored. Scoring them from a transcript would mean laundering one participant's characterization of another into a metric. In a real recorded debate, a narrator asserted “this is where she starts to get upset” — scoring composure from that line would silently ratify one participant's claim about another as measurement.
“Composure: 34” attached to a named real person is a claim that has to be defended. “Declined to continue at 9:27 · evidence” is an observation with a receipt. AgoraCrux ships the second and refuses the first — and says so: coverage reads “4 of 6 signals measured, composure requires audio,” never a silent six-for-six.
Enforced in code today — and checkable, not just claimed.
These aren't values statements. They're constraints on how the system is built, and each one is something you could go verify.
No unearned verdict
A winner is only declared when the room agreed to a rubric before the debate, and only when the evidence actually supports it. Otherwise it's withheld, with the reason stated plainly.
Every score links to its evidence
An index that can't produce the utterances behind it renders as insufficient evidence — not a number.
Numbers are counted, never asked for
The model classifies discrete events with evidence pointers; arithmetic aggregates them. A number a model invents can't be audited, reproduced, or traced. A number counted from stored events can be all three.
Absent, unmeasured, and zero are three different claims
Distinguished everywhere in the interface. A speaker never asked a question shows “—”, not 0%. An unbuilt index shows “·”, not 0. A check that can't run on this format shows “not applicable,” never “passed.”
What's actually shipping, not the whole list.
This is a curated cut. The complete build status, tagged the same honest way, is one click away.
Crux finderLive
Finds where speakers actually diverge versus what they appear to be arguing about — separating the apparent disagreement from the real one. The product is named after it.
Nugget BoardLive
A live tree of typed syntheses — claim, position, pushback, question, evidence — built continuously as people talk, with Scan, Focus, Hide tangents, and a chrono scrub back through the board's own history.
The LabLive
Per-nugget, operator-triggered tools: fact-check, academic fact-check, Steelman, Toulmin analysis, cross-exam questions, and fallacy check against a closed 24-item taxonomy — split formal/informal, every finding paired with its own steelman, every finding dismissible.
Attribution repairLive
Diarization errors get flagged, voiceprint analysis scores whether the audio actually matches the assigned speaker, and an operator listens and approves. Every batch is reversible.
Cost tracking & HTML exportLive
Every Claude and Deepgram call is logged with token counts and cost. The finished session exports as one self-contained HTML file — scorecard, transcript, and the idea board with your layout — zero external requests, openable forever.
AgoraCrux rooms run 5–7 speakers, sometimes more — recorded sessions have already hit 9 and 10. Not a two-person podium.
See the full build status, including what's rough and what's still just a plan
Isn't this just AI doing the debate?
Debate culture has built-in resistance to AI assistance, for good reason — getting caught using it mid-competition reads as cheating. The answer is in what the architecture refuses to do.
It doesn't argue for anyone
It never generates a rebuttal, never scores in secret, never intervenes in the exchange.
It shows its work
It records and organizes what people actually said, and every finding is anchored to the transcript behind it.
Every finding is dismissible
By a human, always. Dismissals persist and drop out of every downstream count.
The room sets the terms
When a verdict is declared at all, it's computed off weights the room agreed to before the debate — the tool does the arithmetic, not the judging.
This isn't AI doing the debate. It's a referee finally holding a stopwatch and a rulebook.
Agora and crux.
Agora — the public square where argument was invented. Crux — the actual point of disagreement, and the name of the feature that finds it. Put together, that's the whole bet: give people back a real square to argue in, and a way to see what they're actually disagreeing about once they're standing in it.
Upload it or link it — the breakdown runs the same as a live room.
Podcasters, streamers, and anyone hosting debates are a primary audience, not an afterthought. Uploading a file or pasting a link puts a recording through the same pipeline a live room gets.
Ingest anythingLive
Upload an audio or video file, or paste a link — the server fetches it and runs it through Deepgram the same as if it happened live in front of the bot.
The same tools, after the factLive
Nugget Board, Crux finder, the Lab, attribution repair — all of it runs on a recording the same as a live capture.
One score doesn't carry over yetPartial
Question tracking runs live, so an uploaded recording currently has no Responsiveness score — it shows “not measured,” never a zero. Diarization accuracy also degrades faster as speaker count rises on uploads; live capture is unaffected, because each speaker has their own channel.
A room your audience joinsVision
Opening a breakdown so a community can follow along at their own pace — pausing on a nugget, checking a fact — is direction, not something shipped yet.
The founding record, lightly proofread.
The sections above are the working translation. This is the original record of how the idea started — proofread for spelling, wording preserved, kept on the page while it's still being built out.
Terms: a nugget is a single synthesized thought. The nugget board is the running collection of nuggets for a debate.
Understanding, not talking past each other
I want people to be able to have a debate/discussion/discourse and not talk past each other. I want them to genuinely be able to understand what the other side is trying to convey — like a steelman version of an argument. This is shown with the synthesis of the thought nuggets on the nugget board.
Real-time fact checking
No more “OK, we're in the middle of a conversation, but can you provide a source for that? I'll wait.” Things like that derail conversations, and when sources are provided it isn't feasible to have someone digest what they're given within the confines of a conversation. Having the ability to pull and synthesize the related information means the conversation can operate on truth and sound arguments.
Finding the crux
People talk past each other, so having the ability to understand the crux of the current back-and-forth is important. Understanding what someone is saying — but not only that, understanding what they are truly trying to convey with their words, and where their words are coming from within the framework of their worldview.
Catching dodges and dog whistles
Actively see if someone is dodging a question, or hearing a dog whistle and fighting immediately with what they perceive someone is saying instead of being walked down a dialogue tree. Sometimes we just need a response to a question directly, so the person who is trying to lay the groundwork of a point can build up to the greater point, making sure the opposition is with them along the way.
Viewer interaction
The ability for viewers in the debate to interact with the program and find definitions of words, or even entire ideas or whole nuggets. Sometimes people talk at such a high level that others don't know what “ontologically” means; or when someone brings up a dense point, the viewers who are not as astute can follow along with a steelmanned, boiled-down understanding that a layman could grasp.
End-of-debate scorecard
A scorecard that shows points conceded, questions left unanswered, and how well each participant did in the context of engagement — things like:
- conceding to another person when appropriate
- steelmanning arguments
- not waffling or wasting talk time
- being good faith and not trying to paint what the other person said in a bad light when it was clear what they were trying to convey
The scorecard will have other features to be determined, but it will be a rich, end-of-debate takeaway that allows each speaker to digest what happened in that exchange — to maybe show how they messed up, or how the other side succeeded in an optics fashion.
The scorecard will also have attached to it:
- the entire transcript, searchable
- the nugget board (searchable and clickable, advancing next and previous to show the flow of the conversation)
Where the idea came from
My initial idea for this app came when I was moderating a debate. I noticed how bad faith people were being in the conversation — intellectually talking above other people and claiming it was because they knew about a subject so much more than the opposition. I say if that is the case, you should be able to synthesize what you are saying into a form that is easily conveyed and digestible.
The program will have something similar to Jackbox TV room codes, so when someone who pays for the program starts up an instance — or a room, so to speak — viewers can join it in a browser. As the debate is going, the viewer can select a concept spoken about in a nugget, or a word, and find an explanation, an example, or how it pertains to the current topic. Something like that will be looked up or explained for them, either by a Google search or by an LLM lookup or something — that part is not fleshed out yet.
Bad after-event recall
Another aspect is the idea that people just have bad after-event recall of how something unfolded. They will think that they outperformed someone when, if you have access to the transcript and the ability to look at a conversation in its entirety, you can really determine who was being good faith, and who was seeking to have a genuine interaction for the sake of exploring their own worldview — ironing it out and reconstructing it based on the input from the debate opponent.
Serving up the scorecard to not only the participants but the viewers means the conversation can be gone over, so they can learn from their missteps or their good moves.
The content creator package
This all falls into something larger for the vision that I would call the content creator package. Someone like Destiny the streamer would have the ability to ingest a debate — uploaded or recorded from online — pass it through, and go over it with their community. The community could join the interactive room so they could follow along at their own pace.
This would work for political debates, online debate panels such as “The Crucible,” and others. There are endless debates recorded on YouTube or other sites, but the only way to engage with them is to watch them and talk about them. Imagine if a content creator could run them through this product and get such a richer, deeper extrapolation from it.
This page now pulls from agoracrux-master-vision.md — 16 recorded founding sessions plus a direct inventory of the codebase. That document is the authoritative source going forward; the entries above stay as the original, proofread record of how the idea started.
See what's actually built.
The full inventory, tagged honestly — what's live, what's rough, what's designed but not shipped, and what's still just the plan.
Be in the room when it opens.
AgoraCrux isn’t open yet. Leave an email and a Discord handle and you’ll hear from us when there’s something real to try — the first live rooms will be small, and this is the list they come from.
Two fields, no newsletter, no forwarding to anyone else. We use the Discord handle because that’s where the early rooms will run.
You’re on the list.
We’ll email you when the first rooms open. Nothing else in the meantime.