COMPARE / 2–4 个模型
模型并排对比
横轴是价格、纵轴是能力 —— 越靠左上,性价比越高。 红色虚线是当前榜单里的性价比前沿线(沿着"更贵但更强"的边界走), 落在它下方的模型意味着"花同样的钱能买到更强的"。
为「模型 4」选择 · Mistral AI
← 返回厂商列表- Devstral 2加入对比 →
- Devstral Medium加入对比 →
- Devstral Small加入对比 →
- Devstral Small 2加入对比 →
- Magistral Medium 1加入对比 →
- Magistral Medium 1.2加入对比 →
- Magistral Small 1加入对比 →
- Magistral Small 1.2加入对比 →
- magistral-medium-2506加入对比 →
- Mistral Large加入对比 →
- Mistral Large 2加入对比 →
- Mistral Large 3加入对比 →
- Mistral Medium加入对比 →
- Mistral Medium 3加入对比 →
- Mistral Medium 3.1加入对比 →
- Mistral Medium 3.5加入对比 →
- Mistral Saba加入对比 →
- Mistral Small加入对比 →
- Mistral Small 3加入对比 →
- Mistral Small 3.1加入对比 →
- Mistral Small 3.2加入对比 →
- Mistral Small 4加入对比 →
- mistral-large-2402加入对比 →
- mistral-large-2411加入对比 →
- mistral-medium-2505加入对比 →
- mistral-medium-2508加入对比 →
- mistral-small-2506加入对比 →
- mistralai/codestral-2508加入对比 →
- mistralai/codestral-embed-2505加入对比 →
- mistralai/devstral-2512加入对比 →
- mistralai/ministral-14b-2512加入对比 →
- mistralai/ministral-3b-2512加入对比 →
- mistralai/ministral-8b-2512加入对比 →
- mistralai/mistral-embed-2312加入对比 →
- mistralai/mistral-large-2407加入对比 →
- mistralai/mistral-large-2512加入对比 →
- mistralai/mistral-nemo加入对比 →
- mistralai/mistral-saba-2502加入对比 →
- mistralai/mistral-small-24b-instruct-2501加入对比 →
- mistralai/mistral-small-2603加入对比 →
- mistralai/mistral-small-3.1-24b-instruct-2503加入对比 →
- mistralai/mistral-small-3.2-24b-instruct-2506加入对比 →
- mistralai/voxtral-mini-transcribe-2602加入对比 →
- mistralai/voxtral-mini-tts-2603加入对比 →
- mistralai/voxtral-small-24b-2507加入对比 →
- Mixtral 8x22B Instruct加入对比 →
- Mixtral 8x7B Instruct加入对比 →
- mixtral-8x22b-instruct-v0.1加入对比 →
PRICE / CAPABILITY FRONTIER
模型斩杀线坐标
215 个模型有完整价格与分值;横轴为对数刻度,标出的点是你选中的模型
图上的点可以直接点:点一下会把它设为「聚焦查看」——在图下标出它的名字与读数, 不占用上方 4 个对比名额(键盘 Tab 也能选中)。再点一次相同的点即可取消聚焦。
币种口径:价格已折算为 CNY(1 USD = 6.724641 CNY,来源 open.er-api.com · 观察于 2026-09-26);口径为每日快照(不追实时),原始口径不受影响。
口径:横轴 = 标准输入价(每百万 tokens,已折人民币),纵轴 = AA 能力指数。 只画同时有价格与分值的行 —— 缺一个就没法定位,硬塞进图里等于编数据。
PROFILE / 能力画像
六轴对照
同一固定区间归一化,跨模型可比;缺数据的轴标"无数据"
athene-70b-0725
athene-70b-0725
这个口径下没有数据(该榜源没有这类指标)
- 能力指数
- 1265.143
- 性价比
- 无数据
- 吞吐
- 无数据
- 上下文
- 无数据
- 输入成本
- 无数据
- 输出成本
- 无数据
Hy3-preview
hy3-preview
- 能力指数
- 22.7362
- 性价比
- 53.67
- 吞吐
- 无数据
- 上下文
- 256000
- 输入成本
- 0.423652383
- 输出成本
- 1.4121746099999999
Solar Pro 4
solar-pro4
- 能力指数
- 28.1541
- 性价比
- 13.96
- 吞吐
- 90.6762
- 上下文
- 512000
- 输入成本
- 2.0173923
- 输出成本
- 8.0695692
MATRIX / 完整对照
各项指标并排对比
24 / 32 项有数据 · 绿色高亮 = 该行最优;"—" 表示该模型这项没有数据(不参与最优比较)
| 指标 | athene-70b-0725 | Hy3-preview | Solar Pro 4 |
|---|---|---|---|
| 综合智能指数 越高越好 | 1,265.14 最优 | 22.74 | 28.15 |
|
编程与工程
scicode / terminalBench 取均值 越高越好 |
— | — | 34.1 |
|
智能体与工具
tau / apex / analyst 取均值 越高越好 |
— | 92.7 最优 | 23.3 |
|
知识与推理
omniscience / gpqa / hle 取均值 越高越好 |
— | 47.5 最优 | 45.8 |
| SciCode 越高越好 | — | — | 44.6% |
| Terminal-Bench 2.1 越高越好 | — | — | 57.3% |
| Terminal-Bench 4.0 越高越好 | — | — | 0.5% |
| τ²-Banking(工具调用) 越高越好 | — | — | 23.3% |
| GPQA(科学问答) 越高越好 | — | 86.7% | 89.1% 最优 |
| HLE(Humanity's Last Exam) 越高越好 | — | 27.8% | 29.2% 最优 |
| 知识准确性 越高越好 | — | 27.9% 最优 | 18.9% |
| 非幻觉率 越高越好 | — | 12.6% | 75.6% 最优 |
| IFBench(指令遵循) 越高越好 | — | 63.1% | — |
| MMMU-Pro(多模态) 越高越好 | — | — | — |
| ITBench-SRE(运维) 越高越好 | — | — | — |
| GDPval(经济价值任务) 越高越好 | — | — | — |
| Analyst Agent(分析师智能体) 越高越好 | — | — | — |
| Apex Agents 越高越好 | — | — | — |
| 指标 | athene-70b-0725 | Hy3-preview | Solar Pro 4 |
|---|---|---|---|
| 输出吞吐 tokens/s 越高越好 | — | — | 90.7 |
| 首 token 延迟 s 越低越好 | — | — | 1.91 |
| 端到端响应 s 越低越好 | — | — | 29.48 |
| 上下文窗口 tokens 越高越好 | — | 256,000 | 512,000 最优 |
| 指标 | athene-70b-0725 | Hy3-preview | Solar Pro 4 |
|---|---|---|---|
| 输入价 ¥/百万 越低越好 | — | 0.4237 最优 | 2.0174 |
| 输出价 ¥/百万 越低越好 | — | 1.4122 最优 | 8.0696 |
| 每任务成本 ¥ 越低越好 | — | — | — |
| 性价比(指数 ÷ 输入价) 分/元 越高越好 | — | 53.7 最优 | 14.0 |
| 指标 | athene-70b-0725 | Hy3-preview | Solar Pro 4 |
|---|---|---|---|
| 厂商 | 未知 | 未知 | 未知 |
| 榜单排名 越低越好 | 278 | 175 | 110 最优 |
| 推理档位 | — | 是 | 是 |
| 开放权重 | — | — | — |
| 模态 | |||
| 发布时间 | — | — | — |
数据来源:能力与评测 / 性能与延迟 / 价格来自第三方榜单当前有效批次 (Artificial Analysis,随批次更新);档案来自本站模型目录。 复合分(编程/智能体/知识)由该组内评测项**取均值**得出(缺项不参与,不按 0 计)—— 口径见方法学。