Results
Models are listed in Artificial Analysis Intelligence Index order, as checked on . That order is just a starting point; it never changes a community score.
| # | Model | Score | Ratings | Spread of 1–5 ratings | Rate |
|---|
Community vs benchmarks
Do benchmark scores match what people experience? For each task type, we line up the community ranking against independent benchmarks. Benchmarks never change a community score.
Loading benchmark results…
How it works
Benchmax is a simple, open survey. Here’s what goes into a score and what to keep in mind when you read one.
- The score is an average
- Add up a model’s current ratings and divide by how many there are. Ratings run from 1 to 5, and nothing else is weighted in.
- Ranks need 10 ratings
- A model gets a rank once it has 10 ratings for the selected task type. Before that, its average is shown as provisional. Models with the same average share a rank.
- One rating per model, per browser
- Rating a model again replaces your earlier rating. Identity isn’t verified, and someone could vote again from another browser. Limits per browser and per network slow down bursts, but can’t stop every abuse.
- Read scores in context
- Voters choose to take part, so the results aren’t a representative survey. Task type, settings and expectations all shape a rating. Check the rating count and spread, and filter by task, before comparing models.
- How benchmarks are compared
- For each task type we take models with at least 30 ratings of that type and a published benchmark result at their highest listed setting, then compare the two orders with a rank correlation. A model is flagged only when its gap holds up after resampling voters and each benchmark’s published uncertainty 1,000 times.
- Where the default order comes from
- The list starts in Artificial Analysis Intelligence Index order among these models, using each model’s highest listed reasoning setting. Qwen3.8 Flash uses its Flash-Next result, which Alibaba says is served as qwen3.8-flash. Switch to Community score to sort by ratings instead.
- Which models are listed
- The newest available model in each general-purpose family we track, as of . Preview models are labeled. Specialist audio and image models and restricted-access models aren’t included. Ratings cover a model across whatever modes and settings people used.
- What we store
- The model, task type, rating, submission times, a pseudonymous browser ID and, for one day, a scrambled form of your network address for rate limiting. We never ask for prompts, comments, names or email addresses. The public data contains totals only, with no individual IDs.