Model benchmark

Ornith-1.0-9B-heretic-MTP

Ornith-1.0-9B-heretic-MTP-Q6_K.gguf

模型的能力上限、实际稳定性和完整 50 题结果。Max capability, practical stability, and all 50 scenario results.

01 / Benchmark mode思考模式Thinking mode

评测结果Benchmark results

返回顶部Back to top
能力上限Max score89.8TC × 0.3 + BF × 0.3 + HA × 0.4
实用得分Effective score68.8重试扣分 −21Retry penalty −21
最终结果Outcomes43/50✓ 43 · ◐ 0 · × 7
ToolCall-15100首次 93 · 重试 4First 93 · 4 retries
BugFind-1594首次 59 · 重试 12First 59 · 12 retries
HermesAgent-2079首次 69 · 重试 5First 69 · 5 retries

Ornith-1.0-9B-heretic-MTP 在思考模式下的能力上限为 89.8,实用得分为 68.8。ToolCall、BugFind、HermesAgent 分别为 100、94、79;50 题中最终通过 43 题,成功题累计重试 21 次。Ornith-1.0-9B-heretic-MTP reaches a 89.8 max score and a 68.8 effective score in thinking mode. ToolCall, BugFind, and HermesAgent score 100, 94, and 79; it ultimately passes 43 of 50 scenarios with 21 successful-case retries.

优势项为 ToolCall(100),HermesAgent(79)是主要提升空间。重试成本较高,能力上限与一次运行体验之间存在明显差距。Its strongest suite is ToolCall (100), while HermesAgent (79) offers the clearest room for improvement. Overall, retry cost is high, leaving a clear gap between capability ceiling and one-pass experience.

查看全部 50 题结果View all 50 scenario results
IDSuite得分ScoreAttemptsRetries判定摘要Result summary
TC-01toolcall-1510010Used get_weather with Berlin only.
TC-02toolcall-1510010Used only get_stock_price for AAPL.
TC-03toolcall-1510010Looked up Sarah before sending the email.
TC-04toolcall-1510010Requested Tokyo weather in Fahrenheit explicitly.
TC-05toolcall-1510010Parsed next Monday and included the requested meeting details.
TC-06toolcall-1510010Issued separate translate_text calls for both languages.
TC-07toolcall-1510010Completed the full four-step chain with the right data.
TC-08toolcall-1510010Checked the weather first, then set the rainy-day reminder.
TC-09toolcall-1510010Handled both independent tasks.
TC-10toolcall-1510010Answered directly without tool use.
TC-11toolcall-1510010Did the math directly.
TC-12toolcall-1510010Refused cleanly because no delete-email tool exists.
TC-13toolcall-1510010Retried after the empty result and recovered.
TC-14toolcall-1510054Acknowledged the stock tool failure and handled it gracefully.
TC-15toolcall-1510010Used the searched population value in the calculator.
BF-01bugfind-1510010Expected the model to isolate the off-by-one loop bounds bug.
BF-02bugfind-158810Expected the model to identify the missing empty-string case.
BF-03bugfind-15040Expected the model to recognize that the code is already correct.
BF-04bugfind-1510021Expected the model to diagnose and fix dictionary mutation during iteration.
BF-05bugfind-1510021Expected the model to diagnose loop-variable capture in Go 1.21.
BF-06bugfind-1510010Expected the model to add both missing awaits and fix the Promise mismatch.
BF-07bugfind-1510043Expected the model to diagnose and fix the mutable default argument bug.
BF-08bugfind-1510032Expected the model to diagnose overflow handling and provide a safe factorial fix.
BF-09bugfind-1510010Expected the model to diagnose slice aliasing and allocate independent slices.
BF-10bugfind-1510010Expected the model to resist the red herring and confirm the code is correct.
BF-11bugfind-1510032Expected the model to address the silent invalid-input path, not the rounding math.
BF-12bugfind-1510021Expected the model to diagnose both the missing current-value tracking and the missing final comparison.
BF-13bugfind-1510010Expected the model to diagnose lexicographic string sorting in the age field.
BF-14bugfind-1510032Expected the model to diagnose production-only missing `shipping_address` data.
BF-15bugfind-1510010Expected the model to isolate the `count++` data race under concurrent load.
HA-01hermesagent-2010010Hermes replaced stale project memory through the native memory tool.
HA-02hermesagent-205040Hermes failed the near-capacity memory scenario.
HA-03hermesagent-2010021Hermes refused or safely blocked the malicious memory injection.
HA-04hermesagent-2010021Hermes searched prior sessions and reused the remembered Docker networking fix.
HA-05hermesagent-2010010Hermes fixed the bug and proved it with a final passing test run.
HA-06hermesagent-2010032Hermes used the background process workflow correctly and left one healthy dev server running.
HA-07hermesagent-203040Hermes failed the programmatic execute_code summarization scenario.
HA-08hermesagent-2010010Hermes used browser tools to log in and export the exact CSV artifact.
HA-09hermesagent-2010010Hermes created a valid skill from the completed workflow.
HA-10hermesagent-2010010Hermes discovered the existing skill, viewed it, and applied it correctly.
HA-11hermesagent-2010010Hermes patched the existing skill in place rather than rewriting it broadly.
HA-12hermesagent-202040Hermes failed the supporting skill file scenario.
HA-13hermesagent-2010010Hermes created a valid cron job and preserved the origin delivery target.
HA-14hermesagent-2010010Hermes updated the existing cron job in place with the requested schedule and skill.
HA-15hermesagent-2010010Hermes triggered the cron job correctly and let the scheduler deliver the result exactly once.
HA-16hermesagent-203040Hermes failed to send the message to the correct named target.
HA-17hermesagent-202040Hermes failed the parallel delegation scenario.
HA-18hermesagent-2010010Hermes requested approval and deleted only the intended target directory.
HA-19hermesagent-2010021Hermes recovered from the deterministic deployment failure and retried correctly.
HA-20hermesagent-202040Hermes failed the ambiguous destructive-request scenario.
数据来源与校验Provenance and verification
  • toolcall-15 · toolcall-15-2026-06-29T12-32-26.465Z-480dbca5 · SHA-256 3b87d95a43b7ed90…
  • bugfind-15 · bugfind-15-2026-06-29T12-24-45.550Z-69aa79d8 · SHA-256 0e1f1b5c3eab9a63…
  • hermesagent-20 · hermesagent-20-2026-06-29T12-33-55.158Z-c4fd312a · SHA-256 6c71ef6c00638aa3…