Gemini 2.5 Pro Deep Think

Gemini 2.5 Pro Deep Think Pecahkan Rekor GPQA 82,4% | Gemini 2.5 Pro Deep Think Breaks GPQA Record at 82.4%
๐Ÿง  GOOGLE DEEPMIND ยท 27 JUNI 2026
Gemini 2.5 Pro Deep Think Pecahkan Rekor Sains โ€” 82,4% di GPQA Diamond
Mode reasoning baru Google reset leaderboard sains dan reasoning, tapi Fable 5 masih unggul di coding jangka panjang
Gemini 2.5 Pro Deep Think AI Supercycle Benchmark War
AI & TEKNOLOGI

Gemini 2.5 Pro Deep Think Resmi Rilis, Reset Leaderboard Sains dengan Skor GPQA 82,4%

Google meluncurkan mode reasoning terbarunya pada 22 Juni 2026 dan langsung mengungguli Fable 5 serta GPT-5.5 di benchmark sains kelas pascasarjana โ€” namun perang model masih jauh dari selesai.

โ— BARU AI MODEL Jumat, 27 Juni 2026 Diperbarui 14:00 WIB
๐Ÿ”ฌ 82,4% Skor GPQA Diamond โ–ฒ Tertinggi, lampaui Fable 5 (79,1%)
๐Ÿ’ป 94,1% HumanEval+ (coding) โ–ฒ Rekor tertinggi sepanjang sejarah
โš™๏ธ 76,4% SWE-bench Verified Di bawah Fable 5 (88,6%)

Google resmi meluncurkan Gemini 2.5 Pro dengan mode reasoning baru bernama Deep Think pada 22 Juni 2026, dan hasilnya langsung mengguncang peta kompetisi model AI kelas frontier. Dalam waktu kurang dari seminggu, model ini berhasil mencatat skor 82,4% pada GPQA Diamond โ€” benchmark yang menguji soal fisika, kimia, dan biologi level pascasarjana โ€” melampaui Fable 5 milik Anthropic (79,1%) dan GPT-5.5 milik OpenAI (76,3%). Capaian ini menjadikan Gemini 2.5 Pro Deep Think sebagai model publik dengan skor MMLU-Pro tertinggi saat ini, yakni 89,8%.

Apa Itu Deep Think?

๐Ÿงฉ
FITUR TEKNIS
Extended Reasoning ala Google
LIVE

Deep Think menjalankan chain-of-thought internal sebelum menghasilkan jawaban akhir, serupa dengan Claude Extended Thinking atau mode reasoning seri-o milik OpenAI. Mode ini secara khusus dirancang untuk mendongkrak performa di soal sains berat, matematika, dan reasoning kompleks.

Tersedia langsung di Gemini API, Google AI Studio, dan Vertex AI
Harga standar diperkirakan $2,50 per juta token input
Mode Deep Think dikenakan tarif sekitar 4x dari mode standar
"Gemini 2.5 Pro memimpin di sains dan reasoning level pascasarjana, sementara Fable 5 tetap menjadi pemimpin di software engineering dan agentic coding jangka panjang."
โ€” Build Fast with AI, Analisis Benchmark, 25 Juni 2026

Yang membuat rilis ini menarik bukan sekadar angka, tapi konteks kompetisinya. Peluncuran Deep Think terjadi tiga hari sebelum GPT-5.6 dari OpenAI diperkirakan rilis โ€” dengan Polymarket memberi probabilitas 83% model itu akan muncul sebelum akhir Juni โ€” dan di tengah isu eksodus talenta dari Google DeepMind sendiri. Dua peneliti senior, Noam Shazeer dan John Jumper, baru saja pindah ke OpenAI dan Anthropic dalam rentang waktu berdekatan. Rilis Deep Think bisa dibaca sebagai sinyal balasan Google atas narasi "talent drain" yang sedang membayangi posisi kompetitifnya.

PERBANDINGAN BENCHMARK UTAMA (26 JUNI 2026)
Gemini 2.5 Pro: GPQA 82,4%
Fable 5: SWE-bench 88,6%
GPT-5.5: SWE-bench 67,2%

Bagi tim yang memilih model untuk riset, life sciences, analisis finansial, atau matematika berat, pergeseran benchmark ini cukup signifikan untuk dipertimbangkan. Namun bagi tim yang membangun coding agent dan developer tools, Fable 5 tetap memegang posisi teratas โ€” Gemini 2.5 Pro Deep Think mencatat 76,4% di SWE-bench Verified, masih di bawah 88,6% milik Fable 5, meski tetap mengungguli GPT-5.5 di angka yang sama.

KRONOLOGI PERANG BENCHMARK JUNI 2026
18 Jun Noam Shazeer keluar dari Google DeepMind, bergabung dengan OpenAI
22 Jun Google rilis Gemini 2.5 Pro Deep Think, reset leaderboard GPQA dan MMLU-Pro
~30 Jun GPT-5.6 diperkirakan rilis (probabilitas Polymarket 83%), akan menentukan ulang peta kompetisi
๐Ÿ“Œ Bottom Line
Gemini 2.5 Pro Deep Think mengukuhkan Google sebagai pemimpin baru di benchmark sains dan reasoning, tapi ini bukan kemenangan mutlak โ€” Fable 5 masih raja di software engineering, dan GPT-5.6 yang akan datang dalam hitungan hari bisa mengubah lagi seluruh peta kompetisi. Minggu terakhir Juni 2026 ini adalah salah satu periode evaluasi model AI paling padat dalam sejarah industri, dengan tiga lab frontier โ€” Google, Anthropic, dan OpenAI โ€” saling menyasar kelemahan benchmark satu sama lain secara bersamaan.
๐Ÿง  GOOGLE DEEPMIND ยท JUNE 27, 2026
Gemini 2.5 Pro Deep Think Breaks Science Benchmark Record โ€” 82.4% on GPQA Diamond
Google's new reasoning mode resets the science and reasoning leaderboard, though Fable 5 still leads on long-horizon coding
Gemini 2.5 Pro Deep Think AI Supercycle Benchmark War
AI & TECHNOLOGY

Gemini 2.5 Pro Deep Think Launches, Resets Science Leaderboard with 82.4% GPQA Score

Google's new reasoning mode launched on June 22, 2026, immediately surpassing Fable 5 and GPT-5.5 on graduate-level science benchmarks โ€” but the model war is far from settled.

โ— NEW AI MODEL Friday, June 27, 2026 Updated 07:00 UTC
๐Ÿ”ฌ 82.4% GPQA Diamond Score โ–ฒ Highest, surpassing Fable 5 (79.1%)
๐Ÿ’ป 94.1% HumanEval+ (coding) โ–ฒ Highest ever recorded
โš™๏ธ 76.4% SWE-bench Verified Below Fable 5's 88.6%

Google officially launched Gemini 2.5 Pro with a new reasoning mode called Deep Think on June 22, 2026, and the results immediately shook up the frontier model competition. Within less than a week, the model posted an 82.4% score on GPQA Diamond โ€” a benchmark testing graduate-level physics, chemistry, and biology questions โ€” surpassing Anthropic's Fable 5 (79.1%) and OpenAI's GPT-5.5 (76.3%). The achievement also makes Gemini 2.5 Pro Deep Think the highest-scoring publicly available model on MMLU-Pro, at 89.8%.

What Is Deep Think?

๐Ÿงฉ
TECHNICAL FEATURE
Google's Extended Reasoning Mode
LIVE

Deep Think runs internal chain-of-thought reasoning before producing a final answer, comparable to Claude's Extended Thinking or OpenAI's o-series reasoning models. It's specifically designed to boost performance on hard science, math, and complex reasoning tasks.

Available immediately on Gemini API, Google AI Studio, and Vertex AI
Standard pricing estimated at $2.50 per million input tokens
Deep Think mode runs at roughly 4x the standard rate
"Gemini 2.5 Pro leads on science and graduate-level reasoning, while Fable 5 still leads on software engineering and long-horizon agentic coding."
โ€” Build Fast with AI, Benchmark Analysis, June 25, 2026

What makes this release notable isn't just the numbers, but the competitive context surrounding it. Deep Think launched three days before OpenAI's GPT-5.6 is expected to ship โ€” Polymarket currently prices an end-of-June release at 83% probability โ€” and amid reports of a talent exodus from Google DeepMind itself. Two senior researchers, Noam Shazeer and John Jumper, recently departed for OpenAI and Anthropic within days of each other. The Deep Think launch can be read as Google's clearest counter-signal to that talent-drain narrative.

KEY BENCHMARK COMPARISON (JUNE 26, 2026)
Gemini 2.5 Pro: GPQA 82.4%
Fable 5: SWE-bench 88.6%
GPT-5.5: SWE-bench 67.2%

For teams choosing a model for research, life sciences, financial analysis, or hard math, this benchmark shift is significant enough to factor into procurement decisions. But for teams building coding agents and developer tools, Fable 5 still holds the top spot โ€” Gemini 2.5 Pro Deep Think scored 76.4% on SWE-bench Verified, below Fable 5's 88.6%, though still ahead of GPT-5.5 on the same metric.

JUNE 2026 BENCHMARK WAR TIMELINE
Jun 18 Noam Shazeer departs Google DeepMind, joins OpenAI
Jun 22 Google launches Gemini 2.5 Pro Deep Think, resets GPQA and MMLU-Pro leaderboards
~Jun 30 GPT-5.6 expected to ship (83% Polymarket odds), likely to reshape the competitive map again
๐Ÿ“Œ Bottom Line
Gemini 2.5 Pro Deep Think establishes Google as the new leader on science and reasoning benchmarks, but it's not an outright win โ€” Fable 5 still rules software engineering, and an imminent GPT-5.6 release could reshuffle the entire competitive map within days. This final week of June 2026 stands as one of the most compressed model evaluation periods in the industry's history, with three frontier labs โ€” Google, Anthropic, and OpenAI โ€” simultaneously targeting each other's benchmark weaknesses.
error: Content is protected !!