Daniel Marama
ยทDecouple Revenue From Headcount

LLM Council

You ask 1 expert. You get 1 answer.

No second opinion, no ranking, no idea if it was any good.

That is how almost everyone uses AI. 1 chat box. 1 model. You cannot tell if the answer is sharp or just sounds sharp.

You could paste the question into 3 chat apps and compare. Nobody does that twice.

Andrej Karpathy shipped a weekend hack that does it for you. An LLM Council. Every model answers, then ranks the others with the names stripped off.

His note from using it: "Quite often, the models are surprisingly willing to select another LLM's response as superior to their own."

I built my own version into The Cockpit. I call it The Council.

Not a panel of models. A boardroom you convene.

You pick the board that fits the call: General, Technical, Creative, Crisis. Each one seats 5 directors with a named remit, and every seat is a genuinely different model. GPT, Claude, DeepSeek, Qwen, Mistral. Not 1 model wearing 5 masks.

They open blind, cross-examine each other, then label the move: unchanged, conceded, or stronger.

First real question I put to it: build a Buffer alternative into The Cockpit? I wanted a yes. After cross-examination, 3 of the 5 directors said don't build it. Cost of that debate: about 5 cents.

Then I shipped 3 things that matter more than the debate itself.

The chair stopped grading its own homework. A different director scores the confidence now. Scorer and synthesizer are never the same brain.

Dissent got a fate. When the split is 3 to 2, the ruling has to name what would change its mind.

The Buffer call was that split. Forced to face its 2 dissenters, the board flipped its verdict: approve, with conditions. 2 tripwires now sit in the ruling. The build runs past 40 hours: kill it, go back to the off-the-shelf tool. LinkedIn breaks the API inside 6 months: decommission.

And the board remembers. Ask something close to an old ruling and it surfaces that ruling first.

Now look at Monday in your business.

Someone brings an idea. Your room reacts to who said it and how sure they sounded. Nobody scores the reasoning. Nobody can find what you decided last quarter.

Fooled by Randomness, Nassim Taleb: with no record, you cannot separate a good decision from a lucky one. Most rooms keep no record, so every meeting restarts from zero.

3 moves you can install this week:

โ†’ Whoever built the option does not get to score it. Not even you.

โ†’ Write down what would change your mind, before you need it.

โ†’ Log the ruling, the confidence, and who dissented. Search it before you convene again.

None of this needs AI. It needs the review step written down, so it survives the person who invented it, including you.

(Built in Gesspark Code)

3