
Google announced Gemini 3.6 Flash on July 21, 2026, along with two other new models in the Flash family. The update targets developers building production AI agents who need better token efficiency, lower latency, and more reliable performance at scale.
What Gemini 3.6 Flash Brings to the Table
Gemini 3.6 Flash is Google’s new workhorse model for coding, knowledge work, and multimodal tasks. According to the Artificial Analysis Index, it uses 17% fewer output tokens than Gemini 3.5 Flash while delivering better quality across the board.
The pricing sits at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. Google claims this makes it cheaper per task than GPT-5.6 Terra Max, Kimi K3, and Qwen 3.7 Max.
Key Benchmark Improvements
- DeepSWE: 49% accuracy vs. 37% for 3.5 Flash (software engineering tasks)
- MLE Bench: 63.9% vs. 49.7% (machine learning research)
- OSWorld-Verified: 83.0% vs. 78.4% (computer use capabilities)
- GDPval-AA v2: 1421 vs. 1349 (knowledge work and document analysis)
Customers including Figma, Harvey, Hebbia, and JetBrains have already tested 3.6 Flash and report improvements in document parsing, chart analysis, report drafting, and code generation workflows.
Gemini 3.5 Flash-Lite: Speed at Scale
The second model, Gemini 3.5 Flash-Lite, runs at 350 output tokens per second according to Artificial Analysis. It’s priced at $0.30 per 1 million input tokens and $2.50 per 1 million output tokens, making it ideal for high-volume agentic workflows.
Flash-Lite outperforms the older Gemini 3 Flash on several benchmarks:
- SWE-Bench Pro: 54.2% vs. 49.6%
- OSWorld-Verified: 74.0% vs. 65.1%
- Terminal-Bench 2.1: 54% vs. 31%
The model includes configurable thinking levels, from minimal (low latency) to higher levels for complex multi-step tasks. Computer use is now a built-in tool in the Gemini API.
Gemini 3.5 Flash Cyber: Security-Focused Model
The third release, Gemini 3.5 Flash Cyber, is fine-tuned specifically for finding and fixing cybersecurity vulnerabilities. It powers Google’s CodeMender agent, which uses multiple Flash Cyber agents working together to produce security reports.
Google calls this the “first truly competitive AI cyber defense system.” On the CyberGym benchmark, CodeMender with Flash Cyber reaches frontier-level performance at a fraction of the cost of larger models.
Due to the dual-use nature of this technology, Flash Cyber will be available exclusively to governments and trusted partners through CodeMender as a limited-access pilot program.
Gemini 3.5 Pro: Still in Testing
Google also confirmed that Gemini 3.5 Pro is currently testing with partners and will be made broadly available when ready. Meanwhile, the team has started what they call their “most ambitious pre-training run yet” for Gemini 4.
Both 3.6 Flash and 3.5 Flash-Lite are available starting today through the Gemini API and Google AI Studio.
Frequently Asked Questions
How much does Gemini 3.6 Flash cost?
Gemini 3.6 Flash costs $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. This is cheaper than 3.5 Flash on a per-task basis due to the 17% reduction in output token usage.
Is Gemini 3.6 Flash better than GPT-5.6?
According to Google’s claims, 3.6 Flash is cheaper per task than GPT-5.6 Terra Max, Kimi K3, and Qwen 3.7 Max. Independent benchmarks show it performing at or slightly below Grok 4.5, but at five times the speed.
What is Gemini 3.5 Flash Cyber used for?
Flash Cyber is fine-tuned for finding and fixing cybersecurity vulnerabilities. It powers Google’s CodeMender agent and will be available to governments and trusted partners through a limited pilot program.
When will Gemini 3.5 Pro be released?
Gemini 3.5 Pro is currently in testing with partners. Google has not announced a specific release date but says it will be available “as soon as it’s ready.”
Can I use Gemini 3.6 Flash for coding?
Yes. 3.6 Flash is optimized for code generation and software engineering tasks. It scored 49% on the DeepSWE benchmark compared to 37% for 3.5 Flash, showing significant improvement in coding workflows.
