Agoracrux
AI debate analysis

Judge the argument, not the room.

AgoraCrux listens to a debate — live, or a recording — and separates what was actually established from what merely performed well. Every score traces back to the words that earned it, and every gap between the two gets named instead of averaged away.

Evidence-linked
Every score can produce the utterances behind it — or it renders as “insufficient evidence,” never a number with nothing under it.
Counted, not guessed
The model classifies discrete events; arithmetic aggregates them. It's never asked to invent a score.
Absent isn't zero
A speaker never asked a question shows “—”. An index not yet built shows “·”. Neither shows 0%.
The non-negotiable principle

An analyzer that only confirms what you already believed isn't an analyzer.

It's a rationalization engine, and people notice eventually. The credibility of everything on this page rests on one rule.

The scorer must be capable of returning a result the user didn't expect.
  • Speaker 2 dominated the exchange but established less than you'd think.
  • Speaker 1's position was not refuted; it was abandoned.
  • This format does not support a comparative verdict.
Why this exists

Online debate is corrupted by flawed metrics.

Judges — human or crowd — reward popularity, charisma, interruption, and speaking volume rather than logical rigor and evidence quality. It happens everywhere: academic debate, workplace argument, legal argument, every comment section on the internet.

Chaos is the default

One founder estimated roughly 90% of their debate experiences were chaotic rather than structured — people talking over each other, endless scroll, noise drowning substance. Chaos isn't incidental to bad metrics; it produces them, because whoever is loudest and fastest wins by default.

There's no referee

Moderation is absent, inconsistent, or biased, so disputes have no legitimate resolution. People expect a neutral arbiter for every other competitive interaction in society. Online discourse doesn't have one.

The result: the person who argued better and the person who appeared to win are routinely different people, and afterward nobody can prove it either way.

The argument

Three ideas carry this. Everything else is table stakes.

The rest of this page, and the full build status, are proof of engineering. These are why it's built at all.

Two-track scoring

Partial

Every score splits into two tracks, computed completely separately: propositional support — whose position was actually established — and dialectical performance — who discharged their obligations in the exchange. These come apart constantly. A speaker can dominate the room and establish nothing. A speaker can hold the better-grounded case and still have the worse night. Naming that gap is the thing nobody else offers.

Headlines are generated from thresholds on the computed indices, never written freehand — a summarizer would just drift toward whatever the loudest voice asserted. Right now two of the four indices behind it are built, so live output can only speak to responsiveness spread; the rest of the picture is on the way.

Dialectical performance → Propositional support →

Illustrative — the scatter view is designed, not shipped yet.

The confidence veil

Spec’d

Before comparing speakers at all, the system checks whether this was even a comparable exchange: does one party have post-hoc narration the other can't answer? Is the cut non-chronological? Are participants in comparable roles? Is speaking time wildly lopsided? When something fires, a fine diagonal hatch is drawn across the score panel — scores stay legible underneath, but visibly qualified — with a caption naming exactly why.

Nothing else in the interface uses hatching, so its presence alone is information. Three states, not two: no hatch and “all clear” on a clean format, no hatch and “not applicable” where a check simply can't run on this format, or the hatch with the flags named. A clean live debate should never have to look identical to a clean upload it wasn't.

Support
Responsiveness
Room Control

/// SCORES DISCOUNTED 40% ///   ambush format · unopposed narration

Concept illustration — designed, not shipped yet.

What it refuses to measure

Live — as policy

Composure, tone, and affect are not scored. Scoring them from a transcript would mean laundering one participant's characterization of another into a metric. In a real recorded debate, a narrator asserted “this is where she starts to get upset” — scoring composure from that line would silently ratify one participant's claim about another as measurement.

“Composure: 34” attached to a named real person is a claim that has to be defended. “Declined to continue at 9:27 · evidence” is an observation with a receipt. AgoraCrux ships the second and refuses the first — and says so: coverage reads “4 of 6 signals measured, composure requires audio,” never a silent six-for-six.

Composure: 34 → Declined to continue at 9:27 · evidence
Four commitments

Enforced in code today — and checkable, not just claimed.

These aren't values statements. They're constraints on how the system is built, and each one is something you could go verify.

No unearned verdict

A winner is only declared when the room agreed to a rubric before the debate, and only when the evidence actually supports it. Otherwise it's withheld, with the reason stated plainly.

Every score links to its evidence

An index that can't produce the utterances behind it renders as insufficient evidence — not a number.

Numbers are counted, never asked for

The model classifies discrete events with evidence pointers; arithmetic aggregates them. A number a model invents can't be audited, reproduced, or traced. A number counted from stored events can be all three.

Absent, unmeasured, and zero are three different claims

Distinguished everywhere in the interface. A speaker never asked a question shows “—”, not 0%. An unbuilt index shows “·”, not 0. A check that can't run on this format shows “not applicable,” never “passed.”

Live today A curated slice

What's actually shipping, not the whole list.

This is a curated cut. The complete build status, tagged the same honest way, is one click away.

Crux finderLive

Finds where speakers actually diverge versus what they appear to be arguing about — separating the apparent disagreement from the real one. The product is named after it.

Nugget BoardLive

A live tree of typed syntheses — claim, position, pushback, question, evidence — built continuously as people talk, with Scan, Focus, Hide tangents, and a chrono scrub back through the board's own history.

The LabLive

Per-nugget, operator-triggered tools: fact-check, academic fact-check, Steelman, Toulmin analysis, cross-exam questions, and fallacy check against a closed 24-item taxonomy — split formal/informal, every finding paired with its own steelman, every finding dismissible.

Attribution repairLive

Diarization errors get flagged, voiceprint analysis scores whether the audio actually matches the assigned speaker, and an operator listens and approves. Every batch is reversible.

Cost tracking & HTML exportLive

Every Claude and Deepgram call is logged with token counts and cost. The finished session exports as one self-contained HTML file — scorecard, transcript, and the idea board with your layout — zero external requests, openable forever.

A real room
A. Reyes24%
J. Okafor19%
S. Kade16%
R. Voss14%
M. Lindqvist12%
T. Abara9%
P. Novak6%

AgoraCrux rooms run 5–7 speakers, sometimes more — recorded sessions have already hit 9 and 10. Not a two-person podium.

See the full build status, including what's rough and what's still just a plan

The objection

Isn't this just AI doing the debate?

Debate culture has built-in resistance to AI assistance, for good reason — getting caught using it mid-competition reads as cheating. The answer is in what the architecture refuses to do.

It doesn't argue for anyone

It never generates a rebuttal, never scores in secret, never intervenes in the exchange.

It shows its work

It records and organizes what people actually said, and every finding is anchored to the transcript behind it.

Every finding is dismissible

By a human, always. Dismissals persist and drop out of every downstream count.

The room sets the terms

When a verdict is declared at all, it's computed off weights the room agreed to before the debate — the tool does the arithmetic, not the judging.

This isn't AI doing the debate. It's a referee finally holding a stopwatch and a rulebook.

The name

Agora and crux.

Agora — the public square where argument was invented. Crux — the actual point of disagreement, and the name of the feature that finds it. Put together, that's the whole bet: give people back a real square to argue in, and a way to see what they're actually disagreeing about once they're standing in it.

For creators

Upload it or link it — the breakdown runs the same as a live room.

Podcasters, streamers, and anyone hosting debates are a primary audience, not an afterthought. Uploading a file or pasting a link puts a recording through the same pipeline a live room gets.

Ingest anythingLive

Upload an audio or video file, or paste a link — the server fetches it and runs it through Deepgram the same as if it happened live in front of the bot.

The same tools, after the factLive

Nugget Board, Crux finder, the Lab, attribution repair — all of it runs on a recording the same as a live capture.

One score doesn't carry over yetPartial

Question tracking runs live, so an uploaded recording currently has no Responsiveness score — it shows “not measured,” never a zero. Diarization accuracy also degrades faster as speaker count rises on uploads; live capture is unaffected, because each speaker has their own channel.

A room your audience joinsVision

Opening a breakdown so a community can follow along at their own pace — pausing on a nugget, checking a fact — is direction, not something shipped yet.

Founder's notes

The founding record, lightly proofread.

The sections above are the working translation. This is the original record of how the idea started — proofread for spelling, wording preserved, kept on the page while it's still being built out.

Temporary · proofread, wording preserved

Terms: a nugget is a single synthesized thought. The nugget board is the running collection of nuggets for a debate.

Understanding, not talking past each other

I want people to be able to have a debate/discussion/discourse and not talk past each other. I want them to genuinely be able to understand what the other side is trying to convey — like a steelman version of an argument. This is shown with the synthesis of the thought nuggets on the nugget board.

Real-time fact checking

No more “OK, we're in the middle of a conversation, but can you provide a source for that? I'll wait.” Things like that derail conversations, and when sources are provided it isn't feasible to have someone digest what they're given within the confines of a conversation. Having the ability to pull and synthesize the related information means the conversation can operate on truth and sound arguments.

Finding the crux

People talk past each other, so having the ability to understand the crux of the current back-and-forth is important. Understanding what someone is saying — but not only that, understanding what they are truly trying to convey with their words, and where their words are coming from within the framework of their worldview.

Catching dodges and dog whistles

Actively see if someone is dodging a question, or hearing a dog whistle and fighting immediately with what they perceive someone is saying instead of being walked down a dialogue tree. Sometimes we just need a response to a question directly, so the person who is trying to lay the groundwork of a point can build up to the greater point, making sure the opposition is with them along the way.

Viewer interaction

The ability for viewers in the debate to interact with the program and find definitions of words, or even entire ideas or whole nuggets. Sometimes people talk at such a high level that others don't know what “ontologically” means; or when someone brings up a dense point, the viewers who are not as astute can follow along with a steelmanned, boiled-down understanding that a layman could grasp.

End-of-debate scorecard

A scorecard that shows points conceded, questions left unanswered, and how well each participant did in the context of engagement — things like:

  • conceding to another person when appropriate
  • steelmanning arguments
  • not waffling or wasting talk time
  • being good faith and not trying to paint what the other person said in a bad light when it was clear what they were trying to convey

The scorecard will have other features to be determined, but it will be a rich, end-of-debate takeaway that allows each speaker to digest what happened in that exchange — to maybe show how they messed up, or how the other side succeeded in an optics fashion.

The scorecard will also have attached to it:

  • the entire transcript, searchable
  • the nugget board (searchable and clickable, advancing next and previous to show the flow of the conversation)

Where the idea came from

My initial idea for this app came when I was moderating a debate. I noticed how bad faith people were being in the conversation — intellectually talking above other people and claiming it was because they knew about a subject so much more than the opposition. I say if that is the case, you should be able to synthesize what you are saying into a form that is easily conveyed and digestible.

The program will have something similar to Jackbox TV room codes, so when someone who pays for the program starts up an instance — or a room, so to speak — viewers can join it in a browser. As the debate is going, the viewer can select a concept spoken about in a nugget, or a word, and find an explanation, an example, or how it pertains to the current topic. Something like that will be looked up or explained for them, either by a Google search or by an LLM lookup or something — that part is not fleshed out yet.

Bad after-event recall

Another aspect is the idea that people just have bad after-event recall of how something unfolded. They will think that they outperformed someone when, if you have access to the transcript and the ability to look at a conversation in its entirety, you can really determine who was being good faith, and who was seeking to have a genuine interaction for the sake of exploring their own worldview — ironing it out and reconstructing it based on the input from the debate opponent.

Serving up the scorecard to not only the participants but the viewers means the conversation can be gone over, so they can learn from their missteps or their good moves.

The content creator package

This all falls into something larger for the vision that I would call the content creator package. Someone like Destiny the streamer would have the ability to ingest a debate — uploaded or recorded from online — pass it through, and go over it with their community. The community could join the interactive room so they could follow along at their own pace.

This would work for political debates, online debate panels such as “The Crucible,” and others. There are endless debates recorded on YouTube or other sites, but the only way to engage with them is to watch them and talk about them. Imagine if a content creator could run them through this product and get such a richer, deeper extrapolation from it.

This page now pulls from agoracrux-master-vision.md — 16 recorded founding sessions plus a direct inventory of the codebase. That document is the authoritative source going forward; the entries above stay as the original, proofread record of how the idea started.

See what's actually built.

The full inventory, tagged honestly — what's live, what's rough, what's designed but not shipped, and what's still just the plan.