The AI Council sends one question to several top AI models at the same time, shows you each answer the moment it arrives, and then has a referee model compare them claim by claim. Instead of trusting a single chatbot, you see exactly where independent models agree and where they contradict each other. It's the thing no vendor's own chatbot offers: ChatGPT, Claude and Gemini each let you re-ask their own tiers, but none of them puts a rival's answer next to its own. This is a feature of Secret Chat AI, the AI aggregator.
Three sizes: Duet, Trio and Quartet
- Duet — two models + referee, and you choose the pair (ChatGPT + Claude by default, or any other combination of the four). The same referee and the same disagreement table as the big panel: a smaller council, not a lesser answer, just faster and cheaper. A Duet is also the size that runs free within the daily quota, on a lighter pair of models (see what it costs below).
- Trio — three models + referee, and again you pick which three (ChatGPT + Claude + Gemini by default, or any other combination of the four). This is the size to reach for when a pair splits one against one: two models contradicting each other tells you there is doubt, a third tells you which way it leans, and it costs about half of what going to the full bench does.
- Quartet — all four models answer in parallel: ChatGPT, Claude, Gemini and Grok, one from each company. A referee then extracts the factual claims, builds a disagreement table (which model asserts, contradicts, or ignores each claim), writes a short synthesis from the claims at least two models agree on, and flags the rest under "Verify before acting". There's nothing to configure: leaving a model out would be leaving out the one that might have disagreed.
Every model on a panel comes from a different company, and that's the whole point: the value of a council is that independent providers disagree, while two tiers of the same family agreeing is far weaker evidence, since they share training data and blind spots. A Quartet uses all four; a Duet and a Trio let you pick which of them answer.
Checking an answer you already have
A council doesn't have to start with a question. Paste an answer you got somewhere else (from ChatGPT, a colleague, a forum, a newsletter) into a Duet, a Trio or a Quartet and ask the panel to check it. Each model goes through the pasted text and says what it thinks holds up, and the referee then compares those verdicts claim by claim, so a claim two independent models both query is flagged in the table rather than buried in two long replies.
It also works in the middle of a conversation. Ask a single model first, then switch that same chat to a council and ask "is that right?". A council reads the whole conversation, so the answer it's checking is the one already on your screen and there's nothing to copy across. That's often the cheapest way to use it: one model to draft, two to check the parts that matter.
When the question needs the web, every model searches it
A council of four models answering from memory would be four training cutoffs agreeing with each other. So whenever an answer turns on anything current or checkable (news, prices, laws, releases, figures, who holds which job), each model answers with live web search attached: ChatGPT, Claude, Gemini and Grok each run their own provider's search and can open the pages they find, and the referee then compares what they came back with. That's what makes a disagreement worth something: two models that both searched and still contradict each other have found different evidence, not merely remembered different things.
You don't choose this, and you aren't charged for it when it's pointless. Before the panel starts, the question is read once and the whole council is set to match it: rewrite this paragraph, translate this, help me think through a decision. Nothing there is looked up, and the panel answers on its faster, cheaper models. Anything that turns on a fact goes to the stronger, web-searching models. The choice is made once for the whole panel, never per model, because a council is only worth reading when its members are comparable: if one model had searched and another had not, their disagreement would be about the panel rather than about your question. And when it's genuinely unclear, the panel goes to the stronger side: an easy question answered too well costs you a little, a hard one answered too cheaply costs you the point of asking.
Two panels are outside this, and for the same reason: they contain a model that cannot search at all, so there is nothing to decide. The free panel is never grounded (see below), and the cheap DeepSeek + GPT pair is not routed either, because DeepSeek USA has no web search at either of its speeds, so pairing it with a searching model would make the two halves of the panel incomparable, which is exactly what the rule above exists to prevent.
On the free panel (DeepSeek Flash and GPT Luna) that never happens: both are fast and cheap but answer from memory, whatever the question. The DeepSeek model we run exposes no web-search tool for us to attach, and the fast GPT tier is never given OpenAI's search, so there you are comparing what two models remember rather than what they found.
It saves time, it does not cost time
The obvious worry about asking four models is that it takes four times as long. It doesn't, because they're asked at the same time. A council is only as slow as its slowest member, never the sum of the panel, so the wait is roughly one model's wait, not four. Doing the same thing by hand is the slow version: four tabs, the same question pasted four times, four waits one after another.
Each answer is revealed in its own tab the moment that model finishes, so you start reading the fastest one while the others are still being written, and the referee's comparison appears once every answer is in.
The larger saving is the reading, not the waiting. Four independent answers to the same question are four screens of text that mostly agree, and the part worth your attention is the sentence or two where they don't, which is exactly the part that's hard to spot by eye, because each model phrases it differently and buries it in a different paragraph. The referee does that comparison for you: it pulls out the factual claims, lines them up model by model, and gives you a ready conclusion — a short synthesis built only from what at least two models independently agreed on, plus a list of the rest under "Verify before acting". So what you get back isn't more to read than one chatbot's answer. It's less, and it comes with a note of how far to trust each part.
What it costs
A council spends credits like any other message: roughly what asking the same question to each panel model separately would cost, plus a small referee step. There's no separate subscription for it: every credit pack and plan runs councils at the same rate, and the pricing page shows what each pack buys in councils. The Duet is free within the daily free quota, on a lighter panel (DeepSeek Flash + GPT Luna), so you can watch two independent models disagree, and have an answer you were given elsewhere checked, without ever paying. Credits buy the branded panel (ChatGPT, Claude, Gemini and Grok, each searching the web live) and the bigger sizes: the three-model Trio and the four-model Quartet. That live search is a real part of the cost, since every provider charges per query on top of the answer itself.
Each answer is priced separately, next to the model that wrote it, with the referee's step stated on its own line, so a council compares what the models charge as well as what they say. That's a comparison you can't make anywhere else: it's the only place several models answer the same question, and the same question can cost twice as much on one model as on another for an answer no better. In your credit history that spend is filed under ChatGPT, Claude, Gemini and Grok like any other message on those models, rather than hidden behind one "council" line.
What to do with a disagreement
The table isn't the end of the job. It tells you which claims are worth more work, which is the part that is otherwise guesswork. So when the referee leaves something unsettled, the app offers the obvious next step underneath the answer: take those claims, and only those, to ChatGPT Sol Research with live sources. It writes the follow-up question for you out of the disputed points and puts it in the message box. You read it before sending, or close the suggestion and it doesn't come back for that answer. A Sol Research run starts at roughly 600 credits and can cost substantially more depending on the question, so the button never sends the request by itself.
That division is deliberate, and it's also the cheaper way round. A council is for agreement across models; a research run is for depth from one. Sending the whole panel off to research everything would pay four models to go deep on the claims they already agreed about, which is most of them. So the panel finds the doubt and one model resolves it.
Privacy — stated honestly
A council multiplies where your question travels: instead of one provider, your question reaches each model on the panel, and the referee model additionally reads the question together with the panel's answers. Every one of those requests goes out the same way as any other Secret Chat AI message: anonymously, with no profile and no account identity attached. As everywhere in the app, anonymization protects who you are, not what you type: whatever the question contains reaches those providers verbatim, so leave out details that identify you. The finished check, every answer and the disagreement table alike, is stored only on your device, like the rest of your chats, and the Session Privacy Report lists every model that took part in the turn, not just the one whose answer you read.
The honest limit
Agreement between models is evidence, not proof. Models are trained on overlapping data and share blind spots, so independent providers can be confidently wrong together; research on cross-model comparison finds correlated errors even across different architectures. What a council reliably gives you is the opposite signal: when models disagree, you know a claim needs checking before you rely on it, and the table shows you exactly which one. Read the synthesis as "what the panel converged on", not as established fact, and treat nothing here as professional advice — for medical, legal or financial decisions, the council tells you which claims to bring to a professional, not which to act on.
When to use it
- Health, legal, money — anywhere a confidently wrong answer is expensive, and exactly the questions people prefer to ask anonymously. The council is a filter for what to verify with a professional, never a substitute for one.
- Facts you are about to repeat — in a report, an email, an argument. "Two of four models contradict this" is worth knowing before you hit send.
- Checking another AI's work — paste in the answer you were given, or switch a chat you have already had to a council, because the answer you already have is often the thing you actually want verified.
For a quick second opinion on a single message inside an ordinary conversation, threads remain the lightweight alternative: branch the question and switch the model. The council is for when you want the comparison done for you, claim by claim.