Local model intelligence index · 2026

让模型的真实能力,
在本地被看见。
See what models can really do,
right on local hardware.

统一硬件、统一题集,分别记录思考与无思考表现。这里不只给一个总分,而是呈现模型在工具调用、代码调试与智能体任务中的真实可用性。One hardware setup and one test set, with thinking and no-thinking results tracked separately. More than a single score: a practical view of tool use, debugging, and agent tasks.

BenchLocal-Results DeepSeek-V4-Flash Gemma family Nex-N2-Mini Ornith family Qwen family Step-3.7-Flash Ternary-Bonsai-27B Hy3-IQ1_M

非严谨学术评测,结果用于参考与简单对比;速度仅记录测试环境。Not a strict academic benchmark. Results support reference and rough comparison; speed only records the test environment.

当前数据集已校验Current dataset validatedRTX 5070 Ti 16GB · 128GB RAM · llama.cpp AMD Ryzen™ AI Max+ 395 · llama.cpp
16模型Models
24配置Configurations
50场景 / 每轮Scenarios / run
01 / Model spotlight

一眼看懂模型差异Understand a model at a glance

重点数据先出现,测试细节按需展开;动态只用于解释状态变化。Key results lead, details follow on demand, and motion only clarifies state changes.

02 / Leaderboard

模型排行榜Model leaderboard

双模式模型合并展示,筛选后重新计算排名。Dual-mode models stay merged; ranking updates after filtering.

#01

Ornith-1.0-35B-Heretic-MTP-APEX

Ornith-1.0-35B-Heretic-MTP-APEX-I-Compact.gguf

思考Thinking
95实用Eff. 90
TC 100BF 98HA 89
无思考No Thinking
92.2实用Eff. 76.2
TC 97BF 97HA 85
能力上限Max score95实用Effective 90
#02

Hy3-IQ1_M

Hy3-IQ1_M.gguf

默认Default
94.4实用Eff. 91.4
TC 100BF 100HA 86
能力上限Max score94.4实用Effective 91.4
#03

DeepSeek-V4-Flash

deepseek-v4-flash-free via opencode

默认Default
93.9实用Eff. 84.9
TC 100BF 93HA 90
能力上限Max score93.9实用Effective 84.9
#04

Ternary-Bonsai-27B

Ternary-Bonsai-27B-Q2_0.gguf

思考Thinking
93.7实用Eff. 87.7
TC 100BF 95HA 88
无思考No Thinking
91实用Eff. 88
TC 100BF 90HA 85
能力上限Max score93.7实用Effective 88
#05

Ornith-1.0-35B-MTP-APEX

Ornith-1.0-35B-MTP-APEX-I-Compact.gguf

思考Thinking
93.5实用Eff. 75.5
TC 100BF 93HA 89
无思考No Thinking
93.2实用Eff. 85.2
TC 100BF 92HA 89
能力上限Max score93.5实用Effective 85.2
#06

Agents-A1

Agents-A1-APEX-I-Compact.gguf

思考Thinking
91.2实用Eff. 71.2
TC 100BF 88HA 87
无思考No Thinking
93.1实用Eff. 61.1
TC 97BF 100HA 85
能力上限Max score93.1实用Effective 71.2
#07

Qwen3.6-27B

Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved.i1-IQ3_M.gguf

思考Thinking
91.9实用Eff. 73.9
TC 100BF 93HA 85
无思考No Thinking
87.8实用Eff. 84.8
TC 97BF 85HA 83
能力上限Max score91.9实用Effective 84.8
#08

Gemma-4-26B-A4B-it-qat

gemma-4-26B-A4B-it-qat-heretic-UD-Q4_K_XL.gguf

思考Thinking
91.8实用Eff. 83.8
TC 93BF 97HA 87
无思考No Thinking
89.8实用Eff. 80.8
TC 97BF 89HA 85
能力上限Max score91.8实用Effective 83.8
#09

Qwen3.6-35B-A3B-uncensored-MTP

Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-APEX-I-Compact.gguf

思考Thinking
89.5实用Eff. 81.5
TC 97BF 88HA 85
无思考No Thinking
91.5实用Eff. 80.5
TC 97BF 96HA 84
能力上限Max score91.5实用Effective 81.5
#10

Gemma-4-12B-it-heretic-QAT

gemma-4-12B-it-heretic-QAT-Q4_0.gguf

思考Thinking
90.3实用Eff. 85.3
TC 97BF 88HA 87
能力上限Max score90.3实用Effective 85.3
#11

Ornith-1.0-9B-heretic-MTP

Ornith-1.0-9B-heretic-MTP-Q6_K.gguf

思考Thinking
89.8实用Eff. 68.8
TC 100BF 94HA 79
能力上限Max score89.8实用Effective 68.8
#12

Qwen-AgentWorld-35B-A3B-APEX-I-Compact

Qwen-AgentWorld-35B-A3B-APEX-I-Compact.gguf

思考Thinking
88.5实用Eff. 87.5
TC 100BF 87HA 81
能力上限Max score88.5实用Effective 87.5
#13

Step-3.7-Flash-APEX-I-Mini

Step-3.7-Flash-APEX-I-Mini.gguf

思考Thinking
88.5实用Eff. 78.5
TC 100BF 87HA 81
能力上限Max score88.5实用Effective 78.5
#14

QwenPaw-Flash-9B

QwenPaw-Flash-9B-heretic-MTP-Q6_K.gguf

思考Thinking
88.4实用Eff. 64.4
TC 100BF 88HA 80
能力上限Max score88.4实用Effective 64.4
#15

Nex-N2-Mini

Huihui-Nex-N2-mini-abliterated-APEX-I-Compact.gguf

思考Thinking
86.7实用Eff. 59.7
TC 93BF 88HA 81
能力上限Max score86.7实用Effective 59.7
#16

Qwen-AgentWorld-35B-A3B-MTP-Uncensored

Qwen-AgentWorld-35B-A3B-MTP-Uncensored-APEX-I-Compact.gguf

思考Thinking
85.6实用Eff. 81.6
TC 93BF 87HA 79
无思考No Thinking
75实用Eff. 75
TC 83BF 83HA 63
能力上限Max score85.6实用Effective 81.6
03 / Method

同一环境,三个真实任务面One environment, three practical dimensions

能力上限看模型能做到什么,实用得分同时计算到达结果所付出的重试成本。Max score shows what a model can achieve; effective score also accounts for retry cost.

01

ToolCall-15

参数提取、多轮上下文、并行调用等工具使用场景。Tool use across parameter extraction, multi-turn context, and parallel calls.

02

BugFind-15

跨语言代码调试,包含无 Bug 与红鲱鱼陷阱题。Cross-language debugging, including no-bug and red-herring traps.

03

HermesAgent-20

记忆、技能、调度、消息投递等智能体任务。Agent work across memory, skills, scheduling, and message delivery.

能力上限Max scoreTC × 0.3 + BF × 0.3 + HA × 0.4
实用得分Effective score能力上限 − 成功题重试次数Max score − successful-case retries