Gemini 2.5 Pro Deep Think Resmi Rilis, Reset Leaderboard Sains dengan Skor GPQA 82,4%
Google meluncurkan mode reasoning terbarunya pada 22 Juni 2026 dan langsung mengungguli Fable 5 serta GPT-5.5 di benchmark sains kelas pascasarjana โ namun perang model masih jauh dari selesai.
Google resmi meluncurkan Gemini 2.5 Pro dengan mode reasoning baru bernama Deep Think pada 22 Juni 2026, dan hasilnya langsung mengguncang peta kompetisi model AI kelas frontier. Dalam waktu kurang dari seminggu, model ini berhasil mencatat skor 82,4% pada GPQA Diamond โ benchmark yang menguji soal fisika, kimia, dan biologi level pascasarjana โ melampaui Fable 5 milik Anthropic (79,1%) dan GPT-5.5 milik OpenAI (76,3%). Capaian ini menjadikan Gemini 2.5 Pro Deep Think sebagai model publik dengan skor MMLU-Pro tertinggi saat ini, yakni 89,8%.
Apa Itu Deep Think?
Deep Think menjalankan chain-of-thought internal sebelum menghasilkan jawaban akhir, serupa dengan Claude Extended Thinking atau mode reasoning seri-o milik OpenAI. Mode ini secara khusus dirancang untuk mendongkrak performa di soal sains berat, matematika, dan reasoning kompleks.
"Gemini 2.5 Pro memimpin di sains dan reasoning level pascasarjana, sementara Fable 5 tetap menjadi pemimpin di software engineering dan agentic coding jangka panjang."โ Build Fast with AI, Analisis Benchmark, 25 Juni 2026
Yang membuat rilis ini menarik bukan sekadar angka, tapi konteks kompetisinya. Peluncuran Deep Think terjadi tiga hari sebelum GPT-5.6 dari OpenAI diperkirakan rilis โ dengan Polymarket memberi probabilitas 83% model itu akan muncul sebelum akhir Juni โ dan di tengah isu eksodus talenta dari Google DeepMind sendiri. Dua peneliti senior, Noam Shazeer dan John Jumper, baru saja pindah ke OpenAI dan Anthropic dalam rentang waktu berdekatan. Rilis Deep Think bisa dibaca sebagai sinyal balasan Google atas narasi "talent drain" yang sedang membayangi posisi kompetitifnya.
Bagi tim yang memilih model untuk riset, life sciences, analisis finansial, atau matematika berat, pergeseran benchmark ini cukup signifikan untuk dipertimbangkan. Namun bagi tim yang membangun coding agent dan developer tools, Fable 5 tetap memegang posisi teratas โ Gemini 2.5 Pro Deep Think mencatat 76,4% di SWE-bench Verified, masih di bawah 88,6% milik Fable 5, meski tetap mengungguli GPT-5.5 di angka yang sama.
Sumber: Build Fast with AI, Google DeepMind ยท Diperbarui 27 Juni 2026
Gemini 2.5 Pro Deep Think Launches, Resets Science Leaderboard with 82.4% GPQA Score
Google's new reasoning mode launched on June 22, 2026, immediately surpassing Fable 5 and GPT-5.5 on graduate-level science benchmarks โ but the model war is far from settled.
Google officially launched Gemini 2.5 Pro with a new reasoning mode called Deep Think on June 22, 2026, and the results immediately shook up the frontier model competition. Within less than a week, the model posted an 82.4% score on GPQA Diamond โ a benchmark testing graduate-level physics, chemistry, and biology questions โ surpassing Anthropic's Fable 5 (79.1%) and OpenAI's GPT-5.5 (76.3%). The achievement also makes Gemini 2.5 Pro Deep Think the highest-scoring publicly available model on MMLU-Pro, at 89.8%.
What Is Deep Think?
Deep Think runs internal chain-of-thought reasoning before producing a final answer, comparable to Claude's Extended Thinking or OpenAI's o-series reasoning models. It's specifically designed to boost performance on hard science, math, and complex reasoning tasks.
"Gemini 2.5 Pro leads on science and graduate-level reasoning, while Fable 5 still leads on software engineering and long-horizon agentic coding."โ Build Fast with AI, Benchmark Analysis, June 25, 2026
What makes this release notable isn't just the numbers, but the competitive context surrounding it. Deep Think launched three days before OpenAI's GPT-5.6 is expected to ship โ Polymarket currently prices an end-of-June release at 83% probability โ and amid reports of a talent exodus from Google DeepMind itself. Two senior researchers, Noam Shazeer and John Jumper, recently departed for OpenAI and Anthropic within days of each other. The Deep Think launch can be read as Google's clearest counter-signal to that talent-drain narrative.
For teams choosing a model for research, life sciences, financial analysis, or hard math, this benchmark shift is significant enough to factor into procurement decisions. But for teams building coding agents and developer tools, Fable 5 still holds the top spot โ Gemini 2.5 Pro Deep Think scored 76.4% on SWE-bench Verified, below Fable 5's 88.6%, though still ahead of GPT-5.5 on the same metric.
Sources: Build Fast with AI, Google DeepMind ยท Updated June 27, 2026




