BenchLocal evolution · 2026

每一次测试,
都留下可复查的记录。
Every benchmark leaves
a record you can inspect.

记录模型数据、评测流程和网站展示方式的持续演进。Tracking the evolution of model data, benchmark workflows, and the way results are presented.

Hy3-IQ1_M

新增 83.3 GB 的 IQ1_M 极低比特量化模型完整结果。能力上限 94.4、实用得分 91.4,ToolCall 与 BugFind 各 100 分,HermesAgent 86 分;50 题中最终通过 45 题,仅需 3 次重试。Added complete results for the 83.3 GB IQ1_M ultra-low-bit quant model. Max score 94.4, effective 91.4; ToolCall and BugFind both reach 100, HermesAgent 86; 45 of 50 scenarios ultimately pass with only 3 retries.

Temperature 0.6 · Top-p 0.9 · Top-k 40In 50 t/s · Out 12 t/s83.3 GBAMD Ryzen™ AI Max+ 395
查看模型详情View model details

BenchLocal-Results V1.0

重构数据采集、评分、页面生成和发布前校验流程。现在以规范化快照作为唯一数据源,自动合并同一模型的思考与无思考结果,并用能力上限与实用得分同时呈现模型表现。Rebuilt the data collection, scoring, page generation, and pre-release validation workflow. Normalized snapshots are now the single source of truth, thinking and no-thinking results are merged by model, and both max and effective scores are presented.

首页、排行榜、模型详情页和更新日志完成全新视觉升级;模型 LOGO、字体与方法插图全部本地化,旧详情页地址继续通过兼容跳转访问。The home page, leaderboard, model details, and changelog received a complete visual refresh. Model logos, fonts, and method artwork are local, while legacy detail URLs remain available through compatibility redirects.

15 models · 23 configurations8 dual-mode comparisons24 validated HTML pages0 remote runtime assets
查看新版排行榜Explore the new leaderboard

Ternary-Bonsai-27B-Q2_0

新增思考与无思考完整结果。思考版能力上限 93.7、实用得分 87.7;无思考版能力上限 91.0、实用得分 88.0。Added complete thinking and no-thinking results. Thinking reaches 93.7 max and 87.7 effective; no-thinking reaches 91.0 max and 88.0 effective.

Temperature 0.7 · Top-p 0.95 · Top-k 20In 876 t/s · Out 61.3 t/s6.67 GB
查看双模式详情View dual-mode details

Ornith-1.0-35B-Heretic-MTP-APEX

思考和无思考结果合并为同一个模型入口,并保留两套完整的 50 题明细。Merged thinking and no-thinking results under one model entry while preserving both complete 50-scenario tables.

查看模型详情View model details

详情页分数与表格渲染修复Detail score and table rendering fixes

修复大分数字段、正负差值以及完整题目表格的显示问题。Fixed headline scores, signed deltas, and complete result-table rendering.

增加评测免责声明与环境说明Benchmark disclaimer and environment notes

明确结果用于参考和简单对比,并补充速度数据与硬件环境的适用边界。Clarified that results support reference and rough comparison, with scope notes for speed and hardware context.

Qwen-AgentWorld-35B-A3B

加入系列模型的多个变体,并统一它们在首页与详情页中的命名和来源信息。Added multiple series variants with consistent naming and source information across home and detail pages.

更新日志上线Changelog launched

开始集中记录模型测试、数据调整和网站功能迭代。Started tracking model benchmarks, data adjustments, and site improvements in one place.