Published September 3, 2026
Google's Gemini 3.8 Flash is now live on Secret Chat AI. As of today it answers your Gemini Web Search and Research questions, and it takes the Gemini seat on the AI Council. Google announced the model on September 2, 2026, three weeks after Gemini 3.7 Flash, and we switched the day after.
This is a short article, because it is a small change. The model is newer and, by Google's account, better at hard multi-step work. The price Google charges for it is unchanged, for now. The wait is about the same too, and we measured that rather than assuming it. And the Pro tier, the thing we were waiting for last time, still hasn't shipped.
What Is Gemini 3.8 Flash?
Gemini 3.8 Flash is the fourth Flash model Google has released since May: 3.5 in May, 3.6 in July, 3.7 in August, and now 3.8 in September. Google calls it its "most intelligent workhorse model" and says it brings "significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains."
The record in Google's public model catalogue, which we pulled directly rather than reading from a press release, says this:
- A 1,048,576-token context window with up to 65,536 tokens of output. A full million tokens of input, the same as 3.7 Flash: a long contract set, a large codebase, or a whole research thread in one conversation.
- It is a thinking model. The catalogue marks it
thinking: true. It reasons before it answers, and you pay for that reasoning as part of the output. - Google's own "latest Flash" pointer has moved to it. Google publishes moving aliases that always resolve to the newest model in each line. On September 3,
gemini-flash-latestanswered asgemini-3.8-flash.
Google itself offers 3.8 Flash to Google AI Pro and Ultra subscribers in the Gemini app, in AI Mode in Search and in Google Sheets, and to developers through the Gemini API, AI Studio, Android Studio, Antigravity and Gemini Enterprise. On Secret Chat you reach the same API model without a Google account.
The benchmark claims in Google's announcement are about agentic and professional work. On DeepSWE v1.1, a long-horizon software engineering test, Google says 3.8 Flash "outperforms most larger frontier models in autonomously solving complex engineering problems end to end." It scores 54.9% on HLE-Verified, which Google cites as evidence of multi-step reasoning across science, humanities and professional fields. And Google says it beats 3.7 Flash on the Vals Finance Agent V2 and Harvey Legal Agent benchmarks. Those are Google's numbers, not ours. We have no way to re-run them, so we report them as claims.
It Works Harder, Which Means It Thinks a Little Longer
Google's own explanation of the improvement is worth quoting, because it tells you what to expect: "3.8 Flash works harder. On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively."
Further down the same post Google says it in plainer terms: "At times, the model might use more tokens to maximize performance, especially at higher effort levels." It adds that developers who care most about efficiency can lower the effort level or "continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads." So this is not a speed release, and Google does not present it as one.
When we moved to 3.7 Flash, we measured that it reasoned less than 3.6 on identical prompts, which you felt as a slightly shorter wait. So we ran the same kind of test again, because a newer model can just as easily go the other way, and no announcement will tell you which one you're getting.
We sent the same eight short prompts to both models, twice each, and counted the thinking tokens each model spent before answering:
- Gemini 3.7 Flash: 6,827 thinking tokens across the 16 calls, 1,223 tokens of answers, 41.5 seconds in total.
- Gemini 3.8 Flash: 7,075 thinking tokens across the same 16 calls, 1,219 tokens of answers, 47.6 seconds in total.
So 3.8 Flash thought about 4% more than 3.7 on our prompts, wrote answers of the same length, and was not faster. Individual prompts went both ways; a haiku made both models deliberate far more than a maths question did. This is a small sample of short questions, so read it as a direction rather than a precise figure. The direction matches what Google says: this release spends a bit more effort per question, and does not trade it for speed.
For a one-line question that means the reply still arrives in two to four seconds. For a long research question it means the model may take a few more reasoning steps before it commits, which is the behaviour Google is advertising.
What It Costs You
Google launched 3.8 Flash at the same introductory price as 3.7 Flash, so nothing about the swap changes what a Gemini Web Search or Research answer costs in credits. On our sample the extra thinking added roughly three percent to the total output, and thinking is billed as output, so expect an answer to cost about what it did last week, give or take a few percent depending on the question.
One thing to know for later. Google's pricing page marks these rates introductory, through December 31, 2026, and states that its standard rate, double the introductory one, applies from January 1, 2027. Credits on Secret Chat track what the provider actually charges, so if Google ends the promotion as announced, a Gemini Flash answer will cost more in credits from that date. We would rather tell you now than surprise you in January.
The Pro Tier Is Still Missing
Last month we wrote that Google had shipped three Flash models since May while its Pro line had not moved since the start of the year. It is now four Flash models, and the Pro line still has not moved: Google's own changelog dates its last Pro release, Gemini 3.1 Pro Preview, to February 19, 2026.
We asked Google's API again on September 3. gemini-pro-latest, Google's own pointer for "the newest Pro", still resolves to gemini-3.1-pro-preview, released in February 2026 from a build stamped January, and still carrying the word preview six and a half months later. Google's pricing page lists no Gemini 3.5, 3.6, 3.7 or 3.8 Pro. Its September 2 announcement says nothing about one.
So on Secret Chat, Gemini's Research mode continues to answer on the newest Flash rather than on a Pro model, and we would rather say so plainly than imply otherwise. When a genuine new Pro model ships, that slot moves back to it.
A Cyber Variant You Won't Find Here
Google released a second model on the same day: Gemini 3.8 Flash Cyber, tuned for finding and patching software vulnerabilities. It is not a general-purpose model and Google is not offering it generally. Access goes to what Google calls "trusted defenders" through a new Fairwind Program: government authorities, critical infrastructure operators and software maintainers. It is not on Secret Chat, and it is not something we could offer even if we wanted to.
What Runs Where on Secret Chat
Here is exactly what the Gemini model does today:
- Smart Agentic (the default): a router reads your message and picks the handling. The router itself runs on Gemini 3.5 Flash Lite; where it sends you is what changed.
- Web Search: now Gemini 3.8 Flash, with Google's own search grounding attached, so answers use current information rather than training-cutoff knowledge.
- Research: now Gemini 3.8 Flash as well, at the deepest setting, for questions worth waiting on.
- Fast Chat: unchanged, on Gemini 3.5 Flash Lite. It is the quickest tier in the family, and a full thinking model would be the wrong tool for a one-line question.
- Image generation: unchanged. Gemini's image mode runs on Nano Banana, a separate image model that a text upgrade doesn't touch.
- The AI Council's Gemini seat: when a council question needs a mid-tier panel, the Gemini seat is now Gemini 3.8 Flash with search grounding attached. On a light panel it stays on 3.5 Flash Lite.
The council point deserves a sentence, because it is the case this app is actually built around. If one lab can go more than six months without shipping its top tier, then which lab you picked is a worse question than what several of them say about the same problem. The AI Council sends one question to models from different companies at once, shows each answer as it lands, and then has a referee lay out where they disagree, claim by claim. Google's Gemini is one voice in that, not the whole answer, and a new Flash release changes one seat at the table rather than the table.
How to Use Gemini 3.8 Flash on Secret Chat AI
Nothing to install or switch on. It's already live:
- Open the Secret Chat app and select the Gemini model in the model picker.
- Leave it on Smart Agentic and let the router decide, or pick a mode yourself next to the message box: Web Search or Research to get 3.8 Flash, Fast Chat for a quick reply.
Google's claims for this release are about multi-step work, so that is what to aim a prompt at. Give it a task with several dependent steps and ask it to show you the chain:
I am going to describe a decision with several moving parts. Work through it step by step: list every assumption you are making, check each one against what I actually told you, and flag the step where a wrong assumption would change the conclusion the most. Then give me your answer and the one thing I should verify before acting on it. Here is the situation:
Private by Design, Whichever Gemini Answers
The upgrade changes the model, not the privacy. Secret Chat is a private gateway between you and the AI providers: your prompts are forwarded through our proxy, so you need no Google account and your requests are never tied to you on Google's side. We build no profile of you, and no chat is ever associated with your identity. Your history lives locally in your own browser, and a prompt exists on our side only for as long as it takes to fetch your answer; there is no stored chat archive on our servers. You can pay with crypto if you would rather not link a card to your AI use.
That matters with Gemini in particular. As we covered in Does Gemini Read Your Data?, a consumer Google account ties AI activity to the same identity as the rest of your Google life. Reaching Gemini 3.8 Flash through Secret Chat lets you use the same model as a stranger: retention may still apply at the provider, but your query reaches the LLM anonymized, not linked to your email or identity.
One honest caveat, as always. "Anonymously" describes the link, not the words: no account identifier travels with your prompt, but that isn't a claim that the text stops being identifying. Write your own name or your case number into a message and it's all still sitting there in the message. Secret Chat removes you from your queries (it doesn't remove the data from your messages), so redacting identifying details before you send is still your call. Anonymity is also not privilege, not a legal exemption, and not a way to put anything beyond the reach of a court that's entitled to it. For the full picture, see our dedicated Private Gemini page.
Conclusion
On the short questions we can measure ourselves, Gemini 3.8 Flash is a modest, real step in the tier where most Gemini questions get answered: the same million-token window, the same price for now, and slightly more thinking per question rather than less, which is what Google's talk of greater diligence predicts. Google reports larger gains on long agentic and professional work; we cannot re-run those benchmarks, so we report them as Google's claims. It powers Gemini's Web Search and Research modes, and the Gemini seat on the AI Council, as of today.
The Pro tier remains the open question, now more than six months and four Flash releases old. When that changes, our Research mode moves with it. And Google's introductory pricing has a stated end date, December 31, 2026, so expect the cost of a Gemini answer to move in January if Google does what it has announced.
Try it today at Secret Chat AI: the full multi-model lineup, your history in your own browser, and every query reaching the model anonymously.
Frequently Asked Questions
- What is Gemini 3.8 Flash?
Gemini 3.8 Flash is the newest model in Google's fast Gemini tier, announced on September 2, 2026, three weeks after Gemini 3.7 Flash. It is a thinking model with a 1,048,576-token context window and up to 65,536 tokens of output, and Google describes it as its most intelligent workhorse model, with gains over 3.7 Flash in software engineering, agentic tasks and multi-step reasoning.
- How is Gemini 3.8 Flash different from 3.7 Flash?
Google says it is more diligent on complex tasks, taking extra reasoning steps and calling tools iteratively, and it reports better scores on agentic benchmarks such as DeepSWE v1.1 and 54.9% on HLE-Verified. In our own test of identical short prompts, 3.8 Flash spent about 4% more thinking tokens than 3.7 Flash, wrote answers of the same length, and was not faster. The context window and Google's price are unchanged.
- Is Gemini 3.8 Flash faster than 3.7 Flash?
Not in our measurement. Across 16 identical calls per model, 3.8 Flash took 47.6 seconds in total against 41.5 for 3.7 Flash, with individual prompts going both ways. Google's own description says the model works harder rather than faster, and its announcement warns that it "might use more tokens to maximize performance, especially at higher effort levels", so a slightly longer wait on hard questions is expected behaviour, not a fault.
- Does Gemini 3.8 Flash cost more credits on Secret Chat?
Not because of the swap. Google launched it at the same introductory price as 3.7 Flash, so a Gemini Web Search or Research answer costs about what it did before, give or take a few percent for the extra thinking. Google states that the introductory pricing ends on December 31, 2026 and that a higher standard rate applies from January 1, 2027; since credits track what the provider charges, Gemini Flash answers will cost more from that date if Google proceeds as announced.
- Is there a Gemini 3.8 Pro, or any new Pro model?
No. As of September 3, 2026, Google's own gemini-pro-latest alias still resolves to gemini-3.1-pro-preview, released on February 19, 2026, and Google's pricing page lists no Gemini 3.5, 3.6, 3.7 or 3.8 Pro. The Flash line has shipped four times since May while the Pro line has not moved.
- Which Gemini modes on Secret Chat use 3.8 Flash?
Web Search and Research both run on Gemini 3.8 Flash, and the Smart Agentic router sends you to them automatically when your question calls for it. The Gemini seat on a mid-tier AI Council panel also runs 3.8 Flash. Fast Chat, the router itself and the light council tier stay on Gemini 3.5 Flash Lite, and image generation runs on Nano Banana, a separate image model.
- What is Gemini 3.8 Flash Cyber, and can I use it on Secret Chat?
It is a variant Google released the same day, tuned for vulnerability detection and automated patching. Google offers it only to trusted defenders such as government authorities, critical infrastructure operators and software maintainers through its Fairwind Program. It is not a general-purpose model and is not available on Secret Chat.
- Do I need a Google account to use Gemini 3.8 Flash?
No. Secret Chat forwards your prompts through its own proxy, so you use Gemini without a Google account and without your requests being tied to you. No profile is built, no chat is associated with your identity, and your history stays in your browser. Retention may still apply at the provider, but your query reaches the LLM anonymized, not linked to your email or identity. Note that the content of your messages is sent verbatim: only your identity is removed.
- Can Gemini 3.8 Flash search the web on Secret Chat?
Yes. The Web Search and Research modes attach Google's own search grounding, so answers combine the model's reasoning with current information instead of being limited to training-cutoff knowledge.