General Chat & Assistants

Whole subtree0 votes

No votes yet. Be the first voter.

Rating criteria

Your tier is simply your own subjective judgment — that's fine. The following are rough guidelines when voting.

  • S

    RecommendedExecutes tasks effortlessly. Produces the intended output with few prompts and turns. Excellent cost performance. Almost no hallucinations.

  • A

    GoodExperienced users can accomplish most tasks. Eventually produces the intended output. Good cost performance. Few hallucinations.

  • B

    AverageProduces the intended output only after many prompts and turns. Average cost performance. Errors and hallucinations occur.

  • C

    Somewhat flawedFairly likely to never produce the intended output no matter how many turns you take. Frequent errors and hallucinations.

  • Nearly impossibleFails to produce the intended output in most cases.

Compare with the Arena benchmark

arena.ai (formerly LMArena) blind-battle Elo, converted into tiers for reference arena.ai Data: LMArena Leaderboard Dataset (CC BY 4.0)

How to read this tier list (what is Arena?)

Arena (formerly LMArena) is a voting site where people compare two anonymous AI answers side by side and pick the better one. Votes from around the world are aggregated into an Elo rating, the same system used in chess.

Tiers come from the Elo gap to the board leader: within 20 = S, within 35 = A, within 60 = B, below that = C.

In the per-service view, each service is represented by its highest-scoring model.

Child threads

Ask in this categoryNo account needed