AI Benchmarks

เทียบโมเดล AI ด้วย Benchmark จริงAI model benchmarks

คะแนนความฉลาด การเขียนโค้ด การทำงานแบบ agent การค้นเว็บ และงานออกแบบ รวมจาก Artificial Analysis, Design Arena และการทดสอบของ OpenRouter — อัปเดตทุกคืนIntelligence, coding, agentic, web-search and design scores from Artificial Analysis, Design Arena and OpenRouter evals — refreshed nightly.

ข้อมูลล่าสุดLast updated 17 Sept 2026

ดัชนีความฉลาด (Intelligence Index)Intelligence Index

คะแนนรวมจากชุดทดสอบความรู้ การใช้เหตุผล และคณิตศาสตร์ ยิ่งสูงยิ่งเก่ง · ราคาเป็น $ ต่อ 1 ล้าน token (input / output)Composite of knowledge, reasoning and maths evals — higher is better · price in $ per 1M tokens (input / output)

  1. 1 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) $5 / $25 · coding 81.6 · agentic 58 53.4
  2. 2 Qwen3.8 Max · coding 68.9 · agentic 49.9 53.4
  3. 3 GPT-6 Astra (max) $11 / $55 · coding 76.9 · agentic 51.5 52.8
  4. 4 Claude Opus 5 (Adaptive Reasoning, Max Effort) $5.5 / $27.5 · coding 78 · agentic 56.2 50.7
  5. 5 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) $10 / $50 · coding 76.5 · agentic 51 49.7
  6. 6 GPT-5.6 Sol (max) $4 / $20 · coding 77.4 · agentic 50.5 47.1
  7. 7 Qwen3.8 Max (0902) $2 / $6 · coding 76.2 · agentic 56.1 45.4
  8. 8 GLM-5.3 (max) $1.4 / $4.4 · coding 74.8 · agentic 53.4 44.9
  9. 9 Grok 4.6 (high) $4 / $12 · coding 76.8 · agentic 53.4 44.4
  10. 10 Kimi K3 (max) $2.1 / $10.95 · coding 76.2 · agentic 50.6 43.8
  11. 11 GPT-5.6 Terra (max) $4 / $24 · coding 76.7 · agentic 43.7 42.3
  12. 12 Claude Opus 4.8 (Adaptive Reasoning, Max Effort) $10 / $50 · coding 74.3 · agentic 42.6 42
  13. 13 GLM-5.3-Flash $0.15 / $0.5 · coding 71.5 · agentic 51.2 41.9
  14. 14 Gemini 3.8 Flash (high) $1.35 / $6.75 · coding 76.3 · agentic 41.1 41.2
  15. 15 Qwen3.8 2.4T A95B $2 / $6 · coding 71.9 · agentic 50.4 40
  16. 16 Muse Spark 1.2 (xhigh) $1.25 / $4.25 · coding 72.2 · agentic 44 39.8
  17. 17 DeepSeek V4.1 Flash (Reasoning, Max Effort) $0.3 / $1.2 39.5
  18. 18 Gemini 3.7 Flash (high) $1.35 / $6.75 · coding 76.1 · agentic 36.4 39.4
  19. 19 Grok 4.5 (high) $4 / $12 · coding 72.4 · agentic 42.1 39.1
  20. 20 GPT-5.5 (xhigh) $2.5 / $15 · coding 74.9 · agentic 37.3 38.6
  21. 21 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) $2.2 / $11 · coding 71.5 · agentic 44.3 38.4
  22. 22 GPT-5.6 Luna (max) $0.4 / $2.4 · coding 71.4 · agentic 42.7 37.5
  23. 23 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) $1.31 / $3.94 · coding 68.8 · agentic 42.3 36.3
  24. 24 DeepSeek V4 Flash 0731 (Reasoning, Max Effort) $0.07 / $0.34 · coding 69.1 · agentic 41.7 34.5
  25. 25 Muse Spark 1.1 (xhigh) $1.25 / $4.25 · coding 71.3 · agentic 27.5 34.3

Source: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings). · as of 16 Sept 2026

ดัชนีการเขียนโค้ด (Coding Index)Coding Index

คะแนนจากชุดทดสอบการเขียนและแก้โค้ดScores from code-writing and code-editing evals

  1. 1 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) $5 / $25 per 1M 81.6
  2. 2 Claude Opus 5 (Adaptive Reasoning, Max Effort) $5.5 / $27.5 per 1M 78
  3. 3 GPT-5.6 Sol (max) $4 / $20 per 1M 77.4
  4. 4 GPT-6 Astra (max) $11 / $55 per 1M 76.9
  5. 5 Grok 4.6 (high) $4 / $12 per 1M 76.8
  6. 6 GPT-5.6 Terra (max) $4 / $24 per 1M 76.7
  7. 7 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) $10 / $50 per 1M 76.5
  8. 8 Gemini 3.8 Flash (high) $1.35 / $6.75 per 1M 76.3
  9. 9 Qwen3.8 Max (0902) $2 / $6 per 1M 76.2
  10. 10 Kimi K3 (max) $2.1 / $10.95 per 1M 76.2
  11. 11 Gemini 3.7 Flash (high) $1.35 / $6.75 per 1M 76.1
  12. 12 GPT-5.5 (xhigh) $2.5 / $15 per 1M 74.9
  13. 13 GLM-5.3 (max) $1.4 / $4.4 per 1M 74.8
  14. 14 Claude Opus 4.8 (Adaptive Reasoning, Max Effort) $10 / $50 per 1M 74.3
  15. 15 Grok 4.5 (high) $4 / $12 per 1M 72.4
  16. 16 Muse Spark 1.2 (xhigh) $1.25 / $4.25 per 1M 72.2
  17. 17 Qwen3.8 2.4T A95B $2 / $6 per 1M 71.9
  18. 18 GLM-5.3-Flash $0.15 / $0.5 per 1M 71.5
  19. 19 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) $2.2 / $11 per 1M 71.5
  20. 20 GPT-5.6 Luna (max) $0.4 / $2.4 per 1M 71.4
  21. 21 Muse Spark 1.1 (xhigh) $1.25 / $4.25 per 1M 71.3
  22. 22 Gemini 3.5 Flash (high) $2.7 / $16.2 per 1M 70.1
  23. 23 Gemini 3.6 Flash (high) $1.35 / $6.75 per 1M 69.2
  24. 24 DeepSeek V4 Flash 0731 (Reasoning, Max Effort) $0.07 / $0.34 per 1M 69.1
  25. 25 Qwen3.8 Max 68.9

Source: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings).

τ²-Bench Airline

AI ทำหน้าที่พนักงานสายการบิน คุยหลายรอบ เรียก tool และต้องทำตามนโยบายเคร่งครัด · % งานที่ทำสำเร็จMulti-turn customer-service agents calling tools under strict policy · % of tasks solved

  1. 1 Google: Gemini 3.7 Flash $0.077 / task · 48 tasks 80.6%
  2. 2 Anthropic: Claude Fable 5 $0.965 / task · 300 tasks 80.2%
  3. 3 Anthropic: Claude Opus 5 $0.499 / task · 300 tasks 79.6%
  4. 4 Amazon: Nova Micro 1.0 $1.28 / task · 100 tasks 78.7%
  5. 5 Qwen: Qwen3.5 397B A17B $0.099 / task · 1350 tasks 78.5%
  6. 6 Z.ai: GLM 5.3 $0.098 / task · 100 tasks 78.3%
  7. 7 Anthropic: Claude Fable 5.1 $0.693 / task · 50 tasks 78%
  8. 8 DeepSeek: DeepSeek V4 Pro 0813 $0.106 / task · 200 tasks 77.7%
  9. 9 Google: Gemini 3 Flash Preview $0.125 / task · 100 tasks 77.3%
  10. 10 StepFun: Step 3.7 Flash $0.020 / task · 50 tasks 77.3%
  11. 11 Qwen: Qwen3.5-122B-A10B $0.148 / task · 500 tasks 77.1%
  12. 12 Anthropic: Claude Opus 4.5 $0.534 / task · 300 tasks 77%
  13. 13 NVIDIA: Nemotron 3 Ultra $0.100 / task · 150 tasks 76.9%
  14. 14 Qwen: Qwen3.8 27B $0.080 / task · 150 tasks 76.9%
  15. 15 Anthropic: Claude Opus 4.7 $0.411 / task · 300 tasks 76.9%
  16. 16 DeepSeek: DeepSeek V4.1 Flash $0.028 / task · 50 tasks 76.7%
  17. 17 OpenAI: GPT-5.6 Sol Pro $1.86 / task · 50 tasks 76.7%
  18. 18 Qwen: Qwen3.8 2.4T A95B $0.238 / task · 100 tasks 76.7%
  19. 19 Anthropic: Claude Sonnet 5 $0.199 / task · 300 tasks 76.7%
  20. 20 OpenAI: GPT-5.6 Sol $0.211 / task · 100 tasks 76.7%
  21. 21 Anthropic: Claude Opus 4.8 $0.497 / task · 300 tasks 76.2%
  22. 22 Google: Gemma 4 31B $0.016 / task · 996 tasks 76.1%
  23. 23 DeepSeek: DeepSeek V4 Pro 0423 $0.042 / task · 1048 tasks 76%
  24. 24 SpaceXAI: Grok 4.6 $0.351 / task · 50 tasks 76%
  25. 25 Anthropic: Claude Opus 4.6 $0.484 / task · 300 tasks 75.7%

Agentic IndexAgentic Index

คะแนนการทำงานหลายขั้นตอนด้วยตัวเองจาก Artificial AnalysisAutonomous multi-step task score from Artificial Analysis

  1. 1 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) 58
  2. 2 Claude Opus 5 (Adaptive Reasoning, Max Effort) 56.2
  3. 3 Qwen3.8 Max (0902) 56.1
  4. 4 GLM-5.3 (max) 53.4
  5. 5 Grok 4.6 (high) 53.4
  6. 6 GPT-6 Astra (max) 51.5
  7. 7 GLM-5.3-Flash 51.2
  8. 8 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) 51
  9. 9 Kimi K3 (max) 50.6
  10. 10 GPT-5.6 Sol (max) 50.5
  11. 11 Qwen3.8 2.4T A95B 50.4
  12. 12 Qwen3.8 Max 49.9
  13. 13 Qwen3.8 27B (xhigh) 46.5
  14. 14 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) 44.3
  15. 15 Muse Spark 1.2 (xhigh) 44

Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). · Source: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings).

GPQA Diamond

คำถามวิทยาศาสตร์ระดับบัณฑิตศึกษาที่ค้นหาคำตอบตรง ๆ ไม่ได้ ต้องใช้เหตุผล · % ตอบถูกGraduate-level science questions that resist lookup · % correct

  1. 1 Google: Gemini 3.1 Pro Preview $0.202 / question 94.4%
  2. 2 OpenAI: GPT-6 Astra $0.099 / question 94.4%
  3. 3 Google: Gemini 3.7 Flash $0.031 / question 94.3%
  4. 4 OpenAI: GPT-5.5 $0.308 / question 93.8%
  5. 5 OpenAI: GPT-5.6 Sol Pro $0.445 / question 93.8%
  6. 6 SpaceXAI: Grok 4.6 $0.122 / question 93.3%
  7. 7 Google: Gemini 3.6 Flash $0.067 / question 92.8%
  8. 8 Google: Gemini 3.5 Flash $0.142 / question 92.8%
  9. 9 OpenAI: GPT-5.6 Sol $0.071 / question 91.9%
  10. 10 MoonshotAI: Kimi K3 $0.147 / question 91.5%
  11. 11 MiniMax: MiniMax M3 $0.030 / question 90.6%
  12. 12 OpenAI: GPT-5.4 $0.147 / question 90.3%
  13. 13 OpenAI: GPT-5.6 Luna Pro $0.044 / question 90.3%
  14. 14 DeepSeek: DeepSeek V4.1 Flash $0.032 / question 90.2%
  15. 15 Anthropic: Claude Fable 5.1 $0.127 / question 90%
  16. 16 Anthropic: Claude Opus 4.8 $0.200 / question 89.6%
  17. 17 DeepSeek: DeepSeek V4 Pro 0813 $0.125 / question 89.3%
  18. 18 Anthropic: Claude Opus 4.7 $0.221 / question 89.1%
  19. 19 Amazon: Nova Micro 1.0 $0.219 / question 89.1%
  20. 20 OpenAI: GPT-5.6 Terra $0.032 / question 89%
  21. 21 DeepSeek: DeepSeek V4 Flash Vision Exp $0.022 / question 88.4%
  22. 22 Google: Gemini 3 Flash Preview $0.132 / question 88.3%
  23. 23 OpenAI: GPT-5.2 $0.133 / question 87.9%
  24. 24 OpenAI: GPT-5.6 Luna $0.007 / question 87.7%
  25. 25 Anthropic: Claude Opus 5 $0.127 / question 87%

Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). · as of 13 Sept 2026

BrowseComp

หาข้อมูลที่หายากบนเว็บจริง ต้องค้นหลายขั้นHard-to-find facts on the live web

  1. 1 Claude Opus 5 · parallel $2.36 / task · 141s 89.2%
  2. 2 GPT-5.6-Sol · perplexity $0.500 / task · 116s 82.4%
  3. 3 Deepseek V4 Flash 0731 · perplexity $0.080 / task · 135s 77%
  4. 4 GPT-5.6-Luna · perplexity $0.100 / task · 111s 74%

WideSearch

เติมข้อมูลลงตารางทั้งตารางFill an entire table from the web

  1. 1 GPT-5.6-Sol · perplexity $0.830 / task · 136s 84%
  2. 2 GPT-5.6-Luna · perplexity $0.170 / task · 165s 83.4%
  3. 3 Deepseek V4 Flash 0731 · perplexity $0.100 / task · 119s 77.1%
  4. 4 Claude Opus 5 · parallel $1.36 / task · 218s 67.3%

DeepSearchQA

คำตอบเป็นรายการ ต้องค้นให้ครบList answers, scored for completeness

  1. 1 Claude Opus 5 · parallel $0.790 / task · 144s 88.7%
  2. 2 GPT-5.6-Sol · parallel $0.850 / task · 131s 88.4%
  3. 3 GPT-5.6-Luna · parallel $0.110 / task · 100s 87.5%
  4. 4 Deepseek V4 Flash 0731 · parallel $0.090 / task · 99s 81.7%

Humanity's Last Exam

คำถามระดับผู้เชี่ยวชาญ ตอบโดยค้นเว็บExpert questions answered with live search

  1. 1 GPT-5.6-Sol · parallel $0.710 / task · 85s 71%
  2. 2 Claude Opus 5 · parallel $0.410 / task · 125s 67.3%

Source: OpenRouter search evals (openrouter.ai/benchmarks) · as of 17 Aug 2026

Design Arena

คนโหวตเทียบผลงานออกแบบจากโมเดลแบบไม่เห็นชื่อ แล้วคิดคะแนน Elo · เลือกหมวดPeople vote on blind side-by-side design outputs, scored as Elo · pick a category

  1. 1 Muse Spark 1.3 (xhigh) win 59.2% · 24,327 battles 1364
  2. 2 Muse Spark 1.3 Max win 58.4% · 9,294 battles 1363
  3. 3 Kimi K3 win 60.9% · 8,854 battles 1353
  4. 4 Claude Fable 5.1 win 56.3% · 6,331 battles 1322
  5. 5 Muse Spark 1.2 win 53.7% · 23,076 battles 1322
  6. 6 Claude Opus 5 win 56.3% · 12,544 battles 1319
  7. 7 GLM-5.3 win 55.7% · 5,911 battles 1317
  8. 8 Gemini 3.7 Flash win 56.3% · 32,523 battles 1315
  9. 9 Gemini 3.8 Flash win 52.6% · 13,883 battles 1315
  10. 10 Gemini 3.6 Flash win 55.7% · 25,029 battles 1313
  11. 11 Claude Fable 5 win 56.5% · 21,846 battles 1308
  12. 12 GLM 5.2 win 53.3% · 26,465 battles 1305
  13. 13 Grok 4.6 win 52.9% · 28,323 battles 1305
  14. 14 Claude Opus 4.6 win 61.2% · 24,054 battles 1304
  15. 15 Claude Opus 4.6 (Thinking) win 60.2% · 14,106 battles 1299
  1. 1 Muse Spark 1.3 (xhigh) win 59.6% · 24,327 battles 1388
  2. 2 Muse Spark 1.3 Max win 57.6% · 9,294 battles 1373
  3. 3 Kimi K3 win 63% · 8,854 battles 1369
  4. 4 Claude Opus 5 win 61.3% · 12,544 battles 1360
  5. 5 Qwen3.8 Max win 58.6% · 3,529 battles 1351
  6. 6 GLM-5.3 win 59.9% · 5,911 battles 1347
  7. 7 Gemini 3.8 Flash win 53.9% · 13,883 battles 1339
  8. 8 Claude Fable 5.1 win 58.4% · 6,331 battles 1337
  9. 9 GLM-5.3-Flash win 55.5% · 13,426 battles 1337
  10. 10 Claude Fable 5 win 58.1% · 21,846 battles 1331
  11. 11 Muse Spark 1.2 win 53.2% · 23,076 battles 1326
  12. 12 Gemini 3.6 Flash win 55.3% · 25,029 battles 1318
  13. 13 Gemini 3.7 Flash win 51.7% · 32,523 battles 1316
  14. 14 Grok 4.6 win 55.1% · 28,323 battles 1314
  15. 15 GLM 5.2 win 55.7% · 26,465 battles 1309
  1. 1 GLM 5.1 win 67% · 5,939 battles 1366
  2. 2 Claude Fable 5.1 win 63.2% · 6,331 battles 1365
  3. 3 Kimi K3 win 64.1% · 8,854 battles 1365
  4. 4 Muse Spark 1.2 win 61.6% · 23,076 battles 1364
  5. 5 Claude Opus 5 win 61.2% · 12,544 battles 1355
  6. 6 Muse Spark 1.3 (xhigh) win 55.3% · 24,327 battles 1331
  7. 7 Gemini 3.7 Flash win 57.1% · 32,523 battles 1330
  8. 8 Claude Fable 5 win 58.1% · 21,846 battles 1325
  9. 9 Gemini 3.6 Flash win 53.5% · 25,029 battles 1312
  10. 10 GLM 5.2 win 53.8% · 26,465 battles 1312
  11. 11 Grok 4.6 win 51.9% · 28,323 battles 1303
  12. 12 Claude Opus 4.6 win 58.7% · 24,054 battles 1296
  13. 13 Qwen3.7 Plus win 52.3% · 7,612 battles 1295
  14. 14 Claude Sonnet 4.6 win 58.2% · 22,174 battles 1294
  15. 15 Muse Spark 1.1 win 52% · 28,947 battles 1292
  1. 1 Claude Fable 5.1 win 66.9% · 6,331 battles 1413
  2. 2 Kimi K3 win 62.6% · 8,854 battles 1400
  3. 3 GLM-5.3 win 64.1% · 5,911 battles 1375
  4. 4 Muse Spark 1.3 Max win 57.7% · 9,294 battles 1372
  5. 5 Claude Opus 5 win 59.5% · 12,544 battles 1364
  6. 6 Claude Fable 5 win 61.5% · 21,846 battles 1361
  7. 7 Muse Spark 1.3 (xhigh) win 56.1% · 24,327 battles 1358
  8. 8 Gemini 3.7 Flash win 56.2% · 32,523 battles 1352
  9. 9 Qwen3.8 Max win 56% · 3,529 battles 1345
  10. 10 Gemini 3.8 Flash win 54% · 13,883 battles 1340
  11. 11 Muse Spark 1.2 win 51.8% · 23,076 battles 1324
  12. 12 Grok 4.6 win 54.6% · 28,323 battles 1322
  13. 13 GPT-5.5 win 58.8% · 29,177 battles 1320
  14. 14 Claude Sonnet 5 win 54% · 25,252 battles 1313
  15. 15 GLM-5.3-Flash win 47.7% · 13,426 battles 1312
  1. 1 Muse Spark 1.3 Max win 64.6% · 9,294 battles 1431
  2. 2 Kimi K3 win 68.9% · 8,854 battles 1425
  3. 3 Claude Fable 5.1 win 68.7% · 6,331 battles 1424
  4. 4 GLM-5.3 win 67.8% · 5,911 battles 1392
  5. 5 Muse Spark 1.3 (xhigh) win 59.3% · 24,327 battles 1387
  6. 6 Claude Opus 5 win 61.7% · 12,544 battles 1363
  7. 7 Qwen3.8 Max win 58.4% · 3,529 battles 1361
  8. 8 GLM-5.3-Flash win 59.5% · 13,426 battles 1354
  9. 9 Claude Fable 5 win 61.9% · 21,846 battles 1347
  10. 10 Gemini 3.7 Flash win 56.3% · 32,523 battles 1339
  11. 11 GLM 5.1 win 62.6% · 5,939 battles 1336
  12. 12 Muse Spark 1.2 win 54.1% · 23,076 battles 1331
  13. 13 Gemini 3.8 Flash win 51.2% · 13,883 battles 1324
  14. 14 GLM 5.2 win 54% · 26,465 battles 1322
  15. 15 Grok 4.6 win 52.8% · 28,323 battles 1306
  1. 1 Claude Fable 5.1 win 64.4% · 6,331 battles 1352
  2. 2 Claude Opus 5 win 62.1% · 12,544 battles 1350
  3. 3 Kimi K3 win 63% · 8,854 battles 1335
  4. 4 Claude Fable 5 win 64.5% · 21,846 battles 1332
  5. 5 GLM-5.3 win 59.3% · 5,911 battles 1324
  6. 6 GLM-5.3-Flash win 57.4% · 13,426 battles 1315
  7. 7 Muse Spark 1.2 win 56.1% · 23,076 battles 1312
  8. 8 Gemini 3.1 Pro Preview win 67.8% · 45,703 battles 1311
  9. 9 Gemini 3.5 Flash win 60.5% · 40,418 battles 1281
  10. 10 Muse Spark 1.1 win 51.6% · 28,947 battles 1275
  11. 11 Gemini 3 Pro Preview win 72.1% · 13,108 battles 1270
  12. 12 Grok 4.6 win 51.6% · 28,323 battles 1268
  13. 13 GPT-5.5 win 57.5% · 29,177 battles 1262
  14. 14 Claude Opus 4.6 (Thinking) win 60.9% · 14,106 battles 1253
  15. 15 GLM 5.2 win 50.4% · 26,465 battles 1248
  1. 1 Kimi K3 win 64.7% · 8,854 battles 1388
  2. 2 Muse Spark 1.3 Max win 59% · 9,294 battles 1374
  3. 3 Muse Spark 1.3 (xhigh) win 58.3% · 24,327 battles 1365
  4. 4 Claude Fable 5.1 win 59.1% · 6,331 battles 1346
  5. 5 Claude Opus 5 win 58.2% · 12,544 battles 1338
  6. 6 GLM-5.3 win 57.7% · 5,911 battles 1330
  7. 7 Muse Spark 1.2 win 53.8% · 23,076 battles 1326
  8. 8 Claude Fable 5 win 58% · 21,846 battles 1325
  9. 9 Gemini 3.7 Flash win 56.2% · 32,523 battles 1323
  10. 10 Gemini 3.8 Flash win 52.8% · 13,883 battles 1323
  11. 11 GLM 5.2 win 53% · 26,465 battles 1309
  12. 12 Grok 4.6 win 53% · 28,323 battles 1309
  13. 13 Qwen3.8 Max win 54% · 3,529 battles 1309
  14. 14 Gemini 3.6 Flash win 53.7% · 25,029 battles 1305
  15. 15 Claude Opus 4.6 win 61.2% · 24,054 battles 1304
  1. 1 MAI-Image-2.6 win 59.4% · 6,154 battles 1295
  2. 2 Riverflow 2.5 Pro win 64.9% · 7,642 battles 1293
  3. 3 MAI-Image-2.6 Flash win 57% · 13,810 battles 1268
  4. 4 Gemini 3.1 Flash Image Gen 2K (Nano Banana 2) win 62.9% · 48,370 battles 1263
  5. 5 Gemini 3.1 Flash Image Gen (Nano Banana 2) win 61.1% · 49,146 battles 1248
  6. 6 MAI-Image-2.5 win 53.4% · 28,915 battles 1248
  7. 7 Gemini 3 Pro Image Gen 2K (Nano Banana Pro) win 61% · 92,599 battles 1241
  8. 8 FLUX.2 [flex] win 54% · 39,725 battles 1209
  9. 9 GPT-Image-1 Mini win 51.4% · 158,215 battles 1209
  10. 10 GPT-Image-1 win 52.7% · 118,421 battles 1202
  11. 11 FLUX.2 [pro] win 48.8% · 169,103 battles 1192
  12. 12 Gemini 3 Pro Image Preview win 54.2% · 36,284 battles 1181
  13. 13 Gemini 2.5 Flash Image (Nano Banana) win 51.4% · 109,337 battles 1169
  14. 14 Gemini 2.5 Flash Image Gen (Nano Banana) win 51.2% · 33,135 battles 1165
  15. 15 FLUX.2 Klein 4B win 37.9% · 72,394 battles 1059
  1. 1 Riverflow 2.5 Pro win 72.5% · 7,642 battles 1385
  2. 2 MAI-Image-2.6 win 61.5% · 6,154 battles 1318
  3. 3 MAI-Image-2.6 Flash win 61% · 13,810 battles 1303
  4. 4 Gemini 3.1 Flash Image Gen 2K (Nano Banana 2) win 65.1% · 48,370 battles 1279
  5. 5 Gemini 3.1 Flash Image Gen (Nano Banana 2) win 64.4% · 49,146 battles 1271
  6. 6 Gemini 3 Pro Image Gen 2K (Nano Banana Pro) win 62.1% · 92,599 battles 1246
  7. 7 MAI-Image-2.5 win 54% · 28,915 battles 1226
  8. 8 Gemini 3 Pro Image Preview win 60.4% · 36,284 battles 1219
  9. 9 FLUX.2 [flex] win 55.5% · 39,725 battles 1203
  10. 10 FLUX.2 [pro] win 49.1% · 169,103 battles 1198
  11. 11 GPT-Image-1 win 53.9% · 118,421 battles 1194
  12. 12 Gemini 2.5 Flash Image Gen (Nano Banana) win 55.6% · 33,135 battles 1189
  13. 13 GPT-Image-1 Mini win 51% · 158,215 battles 1186
  14. 14 Gemini 2.5 Flash Image (Nano Banana) win 53.4% · 109,337 battles 1181
  15. 15 FLUX.2 Klein 4B Distilled win 38.9% · 71,292 battles 1066

Source: Design Arena (www.designarena.ai) via OpenRouter (openrouter.ai/rankings). · as of 16 Sept 2026

โมเดลเปิดตัวล่าสุดNewest models

โมเดลที่เพิ่งเปิดให้ใช้งานผ่าน OpenRouter พร้อมราคาและ contextRecently released on OpenRouter, with pricing and context length

โมเดลModelเปิดตัวReleasedContextInput $/1MOutput $/1M
Union Alphastealth/union-alpha 16 Sept 2026 262K free free
DeepSeek: DeepSeek Pro Latest~deepseek/deepseek-pro-latest 14 Sept 2026 1.0M $0.58 $1.74
DeepSeek: DeepSeek Flash Latest~deepseek/deepseek-flash-latest 14 Sept 2026 1.0M $0.15 $0.6
Inference.net: Schematron V2 Turboinference-net/schematron-v2-turbo 12 Sept 2026 128K $0.03 $0.15
Inference.net: Schematron V2 Smallinference-net/schematron-v2-small 12 Sept 2026 128K $0.05 $0.23
OpenAI: GPT Astra Latest~openai/gpt-astra-latest 11 Sept 2026 1.1M $10 $50
OpenAI: GPT Sol Latest~openai/gpt-sol-latest 11 Sept 2026 1.1M $2 $10
OpenAI: GPT Terra Latest~openai/gpt-terra-latest 11 Sept 2026 1.1M $2 $12
OpenAI: GPT Luna Latest~openai/gpt-luna-latest 11 Sept 2026 1.1M $0.2 $1.2
Sakana: Fugu Ultra v2sakana/fugu-ultra-v2 11 Sept 2026 1M $5 $30
Sakana: Fugu Maxsakana/fugu-max 11 Sept 2026 1M $2 $6
inclusionAI: Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl 10 Sept 2026 131K $0.06 $0.18
inclusionAI: Ling 3.0 Flash VL (free)inclusionai/ling-3.0-flash-vl:free 10 Sept 2026 262K free free
DeepSeek: DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash 10 Sept 2026 1.0M $0.15 $0.6
Inception: Mercury 2.5inception/mercury-2.5 9 Sept 2026 260K $0.04 $0.15
Nex AGI: Nex-N2.5-Mini (free)nex-agi/nex-n2.5-mini:free 9 Sept 2026 262K free free
Nex AGI: Nex-N2.5-Pro (free)nex-agi/nex-n2.5-pro:free 9 Sept 2026 262K free free
OpenAI: GPT-6 Astraopenai/gpt-6-astra 5 Sept 2026 1.1M $10 $50
OpenAI: GPT-6 Astra (batch)openai/gpt-6-astra:batch 5 Sept 2026 1.1M $5 $25
OpenAI: GPT-6 Astra Proopenai/gpt-6-astra-pro 5 Sept 2026 1.1M $10 $50

Source: OpenRouter models API (openrouter.ai/models) · 444 models

แหล่งข้อมูลและวิธีวัดData sources & methodology

ตัวเลขทั้งหมดดึงผ่าน OpenRouter Benchmarks API และ Models API ทุกคืน (01:00 น.) — เราไม่ได้ปรับแก้คะแนนAll numbers are pulled nightly (01:00 Bangkok) through the OpenRouter Benchmarks & Models APIs — we do not adjust any scores.

Artificial Analysis

บริษัทวิจัยอิสระที่รันชุดทดสอบมาตรฐานกับทุกโมเดลด้วยเงื่อนไขเดียวกัน แล้วสรุปเป็นดัชนี Intelligence, Coding และ Agentic พร้อมราคาต่อ tokenIndependent lab that runs the same standard eval suites on every model and publishes Intelligence, Coding and Agentic indexes plus token pricing.

ใช้ในแท็บUsed in: ฉลาดสุด · เขียนโค้ด · AgentIntelligence · Coding · Agents · 101 models · as of 16 Sept 2026

OpenRouter Evals

OpenRouter รันการทดสอบเองแบบทำซ้ำได้ ทุกคะแนนลิงก์ถึง config ค่าใช้จ่าย และ telemetry: GPQA Diamond (วิทยาศาสตร์), τ²-Bench Airline (agent เรียก tool) และชุดค้นเว็บ BrowseComp / WideSearch / DeepSearchQA / HLEReproducible evals run by OpenRouter, each linked to its config, cost and telemetry: GPQA Diamond, τ²-Bench Airline, and the BrowseComp / WideSearch / DeepSearchQA / HLE search suite.

ใช้ในแท็บUsed in: Agent · GPQA · AI ค้นเว็บWeb search · 268 results · as of 13 Sept 2026

Design Arena

ผู้ใช้จริงโหวตเทียบผลงาน (เว็บไซต์ UI กราฟ เกม 3D โลโก้ ฯลฯ) จากสองโมเดลโดยไม่เห็นชื่อ แล้วคำนวณคะแนน Elo และอัตราชนะReal users vote on blind side-by-side outputs (websites, UI, charts, games, 3D, logos…), aggregated into Elo ratings and win rates.

ใช้ในแท็บUsed in: ออกแบบ UIDesign · 14 categories · as of 16 Sept 2026

OpenRouter Models API

แคตตาล็อกโมเดลที่เปิดให้ใช้งานผ่าน API พร้อมวันเปิดตัว context window และราคา input / output ต่อ tokenCatalogue of models available via API with release date, context window and input / output token pricing.

ใช้ในแท็บUsed in: โมเดลใหม่ & ราคาNew & pricing · 444 models

ข้อมูลภายใต้สัญญาอนุญาต CC BY 4.0 ของ OpenRouter Rankings — แสดงเครดิตแหล่งต้นทางตามที่ API กำหนด · Benchmark วัดเฉพาะงานที่ทดสอบ ควรทดลองกับงานจริงของคุณก่อนเลือกใช้Rankings data is licensed CC BY 4.0 by OpenRouter — credited to each original source as the API specifies · Benchmarks cover only the tasks they test; try models on your own workload before choosing.