Research & Academia
Whole subtree ・ 0 votes
No votes yet. Be the first voter.
Rating criteria
※ Your tier is simply your own subjective judgment — that's fine. The following are rough guidelines when voting.
- S
Recommended — Executes tasks effortlessly. Produces the intended output with few prompts and turns. Excellent cost performance. Almost no hallucinations.
- A
Good — Experienced users can accomplish most tasks. Eventually produces the intended output. Good cost performance. Few hallucinations.
- B
Average — Produces the intended output only after many prompts and turns. Average cost performance. Errors and hallucinations occur.
- C
Somewhat flawed — Fairly likely to never produce the intended output no matter how many turns you take. Frequent errors and hallucinations.
- ✗
Nearly impossible — Fails to produce the intended output in most cases.
External benchmark composite (reference)
Papers, review sites, and hands-on media tests converted to tiers and averaged (as of 2026-08-11)
Simple average of per-source tiers (S=4 … C=1). Treat single-source tiers with caution (source count in parentheses).
Averaged from the subcategory benchmark tables and Arena tier tables, with tiers scored S=4 to C=1 (AI Search & Deep Research (Arena), Math (Arena)).
Sources:Arena (search)Arena (math)