AI Council:
Several Models Answer, We Show Where They Disagree

The AI Council sends one question to several top AI models at the same time, shows you each answer the moment it arrives, and then has a referee model compare them claim by claim. Instead of trusting a single chatbot, you see exactly where independent models agree and where they contradict each other. It's the thing no vendor's own chatbot offers: ChatGPT, Claude and Gemini each let you re-ask their own tiers, but none of them puts a rival's answer next to its own. This is a feature of Secret Chat AI, the AI aggregator.

Three sizes: Duet, Trio and Quartet

  • Duettwo models + referee, and you choose the pair (ChatGPT + Claude by default, or any other combination of the four). The same referee and the same disagreement table as the big panel: a smaller council, not a lesser answer, just faster and cheaper. A Duet is also the size that runs free within the daily quota, on a lighter pair of models (see what it costs below).
  • Triothree models + referee, and again you pick which three (ChatGPT + Claude + Gemini by default, or any other combination of the four). This is the size to reach for when a pair splits one against one: two models contradicting each other tells you there is doubt, a third tells you which way it leans, and it costs about half of what going to the full bench does.
  • Quartetall four models answer in parallel: ChatGPT, Claude, Gemini and Grok, one from each company. A referee then extracts the factual claims, builds a disagreement table (which model asserts, contradicts, or ignores each claim), writes a short synthesis from the claims at least two models agree on, and flags the rest under "Verify before acting". There's nothing to configure: leaving a model out would be leaving out the one that might have disagreed.

Every model on a panel comes from a different company, and that's the whole point: the value of a council is that independent providers disagree, while two tiers of the same family agreeing is far weaker evidence, since they share training data and blind spots. A Quartet uses all four; a Duet and a Trio let you pick which of them answer.

Checking an answer you already have

A council doesn't have to start with a question. Paste an answer you got somewhere else (from ChatGPT, a colleague, a forum, a newsletter) into a Duet, a Trio or a Quartet and ask the panel to check it. Each model goes through the pasted text and says what it thinks holds up, and the referee then compares those verdicts claim by claim, so a claim two independent models both query is flagged in the table rather than buried in two long replies.

It also works in the middle of a conversation. Ask a single model first, then switch that same chat to a council and ask "is that right?". A council reads the whole conversation, so the answer it's checking is the one already on your screen and there's nothing to copy across. That's often the cheapest way to use it: one model to draft, two to check the parts that matter.

When the question needs the web, every model searches it

A council of four models answering from memory would be four training cutoffs agreeing with each other. So whenever an answer turns on anything current or checkable (news, prices, laws, releases, figures, who holds which job), each model answers with live web search attached: ChatGPT, Claude, Gemini and Grok each run their own provider's search and can open the pages they find, and the referee then compares what they came back with. That's what makes a disagreement worth something: two models that both searched and still contradict each other have found different evidence, not merely remembered different things.

You don't choose this, and you aren't charged for it when it's pointless. Before the panel starts, the question is read once and the whole council is set to match it: rewrite this paragraph, translate this, help me think through a decision. Nothing there is looked up, and the panel answers on its faster, cheaper models. Anything that turns on a fact goes to the stronger, web-searching models. The choice is made once for the whole panel, never per model, because a council is only worth reading when its members are comparable: if one model had searched and another had not, their disagreement would be about the panel rather than about your question. And when it's genuinely unclear, the panel goes to the stronger side: an easy question answered too well costs you a little, a hard one answered too cheaply costs you the point of asking.

Two panels are outside this, and for the same reason: they contain a model that cannot search at all, so there is nothing to decide. The free panel is never grounded (see below), and the cheap DeepSeek + GPT pair is not routed either, because DeepSeek USA has no web search at either of its speeds, so pairing it with a searching model would make the two halves of the panel incomparable, which is exactly what the rule above exists to prevent.

On the free panel (DeepSeek Flash and GPT Luna) that never happens: both are fast and cheap but answer from memory, whatever the question. The DeepSeek model we run exposes no web-search tool for us to attach, and the fast GPT tier is never given OpenAI's search, so there you are comparing what two models remember rather than what they found.

It saves time, it does not cost time

The obvious worry about asking four models is that it takes four times as long. It doesn't, because they're asked at the same time. A council is only as slow as its slowest member, never the sum of the panel, so the wait is roughly one model's wait, not four. Doing the same thing by hand is the slow version: four tabs, the same question pasted four times, four waits one after another.

Each answer is revealed in its own tab the moment that model finishes, so you start reading the fastest one while the others are still being written, and the referee's comparison appears once every answer is in.

The larger saving is the reading, not the waiting. Four independent answers to the same question are four screens of text that mostly agree, and the part worth your attention is the sentence or two where they don't, which is exactly the part that's hard to spot by eye, because each model phrases it differently and buries it in a different paragraph. The referee does that comparison for you: it pulls out the factual claims, lines them up model by model, and gives you a ready conclusion — a short synthesis built only from what at least two models independently agreed on, plus a list of the rest under "Verify before acting". So what you get back isn't more to read than one chatbot's answer. It's less, and it comes with a note of how far to trust each part.

What it costs

A council spends credits like any other message: roughly what asking the same question to each panel model separately would cost, plus a small referee step. There's no separate subscription for it: every credit pack and plan runs councils at the same rate, and the pricing page shows what each pack buys in councils. The Duet is free within the daily free quota, on a lighter panel (DeepSeek Flash + GPT Luna), so you can watch two independent models disagree, and have an answer you were given elsewhere checked, without ever paying. Credits buy the branded panel (ChatGPT, Claude, Gemini and Grok, each searching the web live) and the bigger sizes: the three-model Trio and the four-model Quartet. That live search is a real part of the cost, since every provider charges per query on top of the answer itself.

Each answer is priced separately, next to the model that wrote it, with the referee's step stated on its own line, so a council compares what the models charge as well as what they say. That's a comparison you can't make anywhere else: it's the only place several models answer the same question, and the same question can cost twice as much on one model as on another for an answer no better. In your credit history that spend is filed under ChatGPT, Claude, Gemini and Grok like any other message on those models, rather than hidden behind one "council" line.

What to do with a disagreement

The table isn't the end of the job. It tells you which claims are worth more work, which is the part that is otherwise guesswork. So when the referee leaves something unsettled, the app offers the obvious next step underneath the answer: take those claims, and only those, to ChatGPT Sol Research with live sources. It writes the follow-up question for you out of the disputed points and puts it in the message box. You read it before sending, or close the suggestion and it doesn't come back for that answer. A Sol Research run starts at roughly 600 credits and can cost substantially more depending on the question, so the button never sends the request by itself.

That division is deliberate, and it's also the cheaper way round. A council is for agreement across models; a research run is for depth from one. Sending the whole panel off to research everything would pay four models to go deep on the claims they already agreed about, which is most of them. So the panel finds the doubt and one model resolves it.

Privacy — stated honestly

A council multiplies where your question travels: instead of one provider, your question reaches each model on the panel, and the referee model additionally reads the question together with the panel's answers. Every one of those requests goes out the same way as any other Secret Chat AI message: anonymously, with no profile and no account identity attached. As everywhere in the app, anonymization protects who you are, not what you type: whatever the question contains reaches those providers verbatim, so leave out details that identify you. The finished check, every answer and the disagreement table alike, is stored only on your device, like the rest of your chats, and the Session Privacy Report lists every model that took part in the turn, not just the one whose answer you read.

The honest limit

Agreement between models is evidence, not proof. Models are trained on overlapping data and share blind spots, so independent providers can be confidently wrong together; research on cross-model comparison finds correlated errors even across different architectures. What a council reliably gives you is the opposite signal: when models disagree, you know a claim needs checking before you rely on it, and the table shows you exactly which one. Read the synthesis as "what the panel converged on", not as established fact, and treat nothing here as professional advice — for medical, legal or financial decisions, the council tells you which claims to bring to a professional, not which to act on.

When to use it

  • Health, legal, money — anywhere a confidently wrong answer is expensive, and exactly the questions people prefer to ask anonymously. The council is a filter for what to verify with a professional, never a substitute for one.
  • Facts you are about to repeat — in a report, an email, an argument. "Two of four models contradict this" is worth knowing before you hit send.
  • Checking another AI's work — paste in the answer you were given, or switch a chat you have already had to a council, because the answer you already have is often the thing you actually want verified.

For a quick second opinion on a single message inside an ordinary conversation, threads remain the lightweight alternative: branch the question and switch the model. The council is for when you want the comparison done for you, claim by claim.

Frequently Asked Questions

What is the AI Council?
It sends one question to several top AI models at the same time (two, three or four, always one per company), shows you each answer the moment it arrives, and then has a referee model compare them claim by claim. You get a disagreement table showing which model asserts each factual claim, which contradicts it and which never mentions it, a short synthesis of what at least two models agreed on, and a list of what to verify before acting.
Is the AI Council free?
A two-model Duet runs free within the daily quota, on a lighter pair of models that answer from memory rather than searching the web. It's the same feature on cheaper engines, with the same referee and the same disagreement table, so you can see exactly what a council does before paying. Credits buy the branded panels (ChatGPT, Claude, Gemini and Grok, each able to search the web live) and the larger three-model Trio and four-model Quartet.
How is this different from just asking two chatbots the same question?
The comparison is done for you, and it's done on the claims rather than on the vibe. Asking twice leaves you two long replies to reconcile by hand, and the differences that matter are usually a sentence buried in each. The referee extracts the factual claims and puts them side by side, so a contradiction is a row in a table instead of something you have to notice. The models also run in parallel, so a council takes as long as its slowest member rather than the sum of the panel.
Does asking several models make me wait longer?
No. The panel is queried in parallel, so the wait is roughly that of the slowest single model rather than the sum of them. The first answer appears the moment the fastest model finishes, so you're reading while the rest are still being written. Doing the same check by hand is the slow version: four tabs, the same question pasted four times, four waits in a row. The bigger saving is that you don't have to read four answers closely to find the one place they differ; the referee extracts the claims and hands you a short conclusion plus the list of points to verify.
Does the council search the web?
When the question turns on something current or checkable (news, prices, laws, releases, figures), every model on the panel answers with its own provider's live web search attached. That choice is made once for the whole panel rather than per model, because a council is only worth reading when its members are comparable. Questions that need no lookup, such as rewriting or translating a passage, are answered on faster models without search. The free pair never searches.
Can I use it to check an answer I already have?
Yes, and it's one of the most useful ways to run it. Paste an answer you were given anywhere else (another chatbot, a colleague, a newsletter) into a council and ask the panel to check it. You can also switch a conversation you have already had to a council and ask "is that right?": the panel reads the whole conversation, so the answer under review is the one already on your screen.
If the models all agree, is the answer correct?
Not necessarily, and this is the limit worth understanding. Models are trained on overlapping data and share blind spots, so independent providers can be confidently wrong together. Agreement is evidence, not proof. What a council reliably gives you is the opposite signal: when models disagree, you know a claim needs checking before you rely on it, and the table shows you exactly which one.
Is a council as private as an ordinary chat?
Each request is sent the same way as any other message here: anonymously, with no profile and no account identity attached. The finished check is stored only on your device. But a council does widen where your question travels: instead of one provider it reaches every model on the panel, and the referee reads it again alongside their answers. As everywhere in the app, anonymization protects who you are, not what you type, so the advice to leave out identifying details applies with more force here, not less. The Session Privacy Report lists every model that took part.