Ternary-Bonsai-27B
Ternary-Bonsai-27B-Q2_0.gguf · 6.67 GB · Out 61.3 t/s
统一硬件、统一题集,分别记录思考与无思考表现。这里不只给一个总分,而是呈现模型在工具调用、代码调试与智能体任务中的真实可用性。One hardware setup and one test set, with thinking and no-thinking results tracked separately. More than a single score: a practical view of tool use, debugging, and agent tasks.
非严谨学术评测,结果用于参考与简单对比;速度仅记录测试环境。Not a strict academic benchmark. Results support reference and rough comparison; speed only records the test environment.
重点数据先出现,测试细节按需展开;动态只用于解释状态变化。Key results lead, details follow on demand, and motion only clarifies state changes.
Ternary-Bonsai-27B-Q2_0.gguf · 6.67 GB · Out 61.3 t/s
双模式模型合并展示,筛选后重新计算排名。Dual-mode models stay merged; ranking updates after filtering.
Ornith-1.0-35B-Heretic-MTP-APEX-I-Compact.gguf
Hy3-IQ1_M.gguf
deepseek-v4-flash-free via opencode
Ternary-Bonsai-27B-Q2_0.gguf
Ornith-1.0-35B-MTP-APEX-I-Compact.gguf
Agents-A1-APEX-I-Compact.gguf
Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved.i1-IQ3_M.gguf
gemma-4-26B-A4B-it-qat-heretic-UD-Q4_K_XL.gguf
Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-APEX-I-Compact.gguf
gemma-4-12B-it-heretic-QAT-Q4_0.gguf
Ornith-1.0-9B-heretic-MTP-Q6_K.gguf
Qwen-AgentWorld-35B-A3B-APEX-I-Compact.gguf
Step-3.7-Flash-APEX-I-Mini.gguf
QwenPaw-Flash-9B-heretic-MTP-Q6_K.gguf
Huihui-Nex-N2-mini-abliterated-APEX-I-Compact.gguf
Qwen-AgentWorld-35B-A3B-MTP-Uncensored-APEX-I-Compact.gguf
能力上限看模型能做到什么,实用得分同时计算到达结果所付出的重试成本。Max score shows what a model can achieve; effective score also accounts for retry cost.
01参数提取、多轮上下文、并行调用等工具使用场景。Tool use across parameter extraction, multi-turn context, and parallel calls.
02跨语言代码调试,包含无 Bug 与红鲱鱼陷阱题。Cross-language debugging, including no-bug and red-herring traps.
03记忆、技能、调度、消息投递等智能体任务。Agent work across memory, skills, scheduling, and message delivery.