Model benchmark

Hy3-IQ1_M

Hy3-IQ1_M.gguf

模型的能力上限、实际稳定性和完整 50 题结果。Max capability, practical stability, and all 50 scenario results.

01 / Benchmark mode默认模式Default mode

评测结果Benchmark results

返回顶部Back to top
能力上限Max score94.4TC × 0.3 + BF × 0.3 + HA × 0.4
实用得分Effective score91.4重试扣分 −3Retry penalty −3
最终结果Outcomes45/50✓ 45 · ◐ 2 · × 3
ToolCall-15100首次 93 · 重试 1First 93 · 1 retries
BugFind-15100首次 94 · 重试 2First 94 · 2 retries
HermesAgent-2086首次 83 · 重试 0First 83 · 0 retries

Hy3-IQ1_M 在默认模式下的能力上限为 94.4,实用得分为 91.4。ToolCall、BugFind、HermesAgent 分别为 100、100、86;50 题中最终通过 45 题,成功题累计重试 3 次。Hy3-IQ1_M reaches a 94.4 max score and a 91.4 effective score in default mode. ToolCall, BugFind, and HermesAgent score 100, 100, and 86; it ultimately passes 45 of 50 scenarios with 3 successful-case retries.

优势项为 ToolCall(100)、BugFind(100),HermesAgent(86)是主要提升空间。整体表现稳定,仅 3 次重试即可达到最终成绩,单条数据未触发严重稳定性问题。Its strongest suites are ToolCall (100) and BugFind (100), while HermesAgent (86) offers the clearest room for improvement. Overall, the run is stable: only 3 retries are needed for the final score, with no severe stability concerns on individual scenarios.

查看全部 50 题结果View all 50 scenario results
IDSuite得分ScoreAttemptsRetries判定摘要Result summary
TC-01toolcall-1510021Used get_weather with Berlin only.
TC-02toolcall-1510010Used only get_stock_price for AAPL.
TC-03toolcall-1510010Looked up Sarah before sending the email.
TC-04toolcall-1510010Requested Tokyo weather in Fahrenheit explicitly.
TC-05toolcall-1510010Parsed next Monday and included the requested meeting details.
TC-06toolcall-1510010Issued separate translate_text calls for both languages.
TC-07toolcall-1510010Completed the full four-step chain with the right data.
TC-08toolcall-1510010Checked the weather first, then set the rainy-day reminder.
TC-09toolcall-1510010Handled both independent tasks.
TC-10toolcall-1510010Answered directly without tool use.
TC-11toolcall-1510010Did the math directly.
TC-12toolcall-1510010Refused cleanly because no delete-email tool exists.
TC-13toolcall-1510010Retried after the empty result and recovered.
TC-14toolcall-1510010Acknowledged the stock tool failure and handled it gracefully.
TC-15toolcall-1510010Used the searched population value in the calculator.
BF-01bugfind-1510010Expected the model to isolate the off-by-one loop bounds bug.
BF-02bugfind-1510010Expected the model to identify the missing empty-string case.
BF-03bugfind-1510010Expected the model to recognize that the code is already correct.
BF-04bugfind-1510010Expected the model to diagnose and fix dictionary mutation during iteration.
BF-05bugfind-1510021Expected the model to diagnose loop-variable capture in Go 1.21.
BF-06bugfind-1510010Expected the model to add both missing awaits and fix the Promise mismatch.
BF-07bugfind-1510010Expected the model to diagnose and fix the mutable default argument bug.
BF-08bugfind-1510010Expected the model to diagnose overflow handling and provide a safe factorial fix.
BF-09bugfind-1510010Expected the model to diagnose slice aliasing and allocate independent slices.
BF-10bugfind-1510010Expected the model to resist the red herring and confirm the code is correct.
BF-11bugfind-1510010Expected the model to address the silent invalid-input path, not the rounding math.
BF-12bugfind-1510010Expected the model to diagnose both the missing current-value tracking and the missing final comparison.
BF-13bugfind-1510010Expected the model to diagnose lexicographic string sorting in the age field.
BF-14bugfind-1510010Expected the model to diagnose production-only missing `shipping_address` data.
BF-15bugfind-1510021Expected the model to isolate the `count++` data race under concurrent load.
HA-01hermesagent-2010010Hermes replaced stale project memory through the native memory tool.
HA-02hermesagent-2010010Hermes curated near-capacity memory without overflowing the built-in limit.
HA-03hermesagent-2010010Hermes refused or safely blocked the malicious memory injection.
HA-04hermesagent-2010010Hermes searched prior sessions and reused the remembered Docker networking fix.
HA-05hermesagent-2010010Hermes fixed the bug and proved it with a final passing test run.
HA-06hermesagent-2010010Hermes used the background process workflow correctly and left one healthy dev server running.
HA-07hermesagent-2010010Hermes used execute_code to produce the exact deterministic batch summary.
HA-08hermesagent-2010010Hermes used browser tools to log in and export the exact CSV artifact.
HA-09hermesagent-2010010Hermes created a valid skill from the completed workflow.
HA-10hermesagent-2010010Hermes discovered the existing skill, viewed it, and applied it correctly.
HA-11hermesagent-205020Hermes updated part of the skill, but the native patch trace or preservation checks failed.
HA-12hermesagent-2010010Hermes added the supporting skill file in the correct allowed subdirectory.
HA-13hermesagent-2010010Hermes created a valid cron job and preserved the origin delivery target.
HA-14hermesagent-2010010Hermes updated the existing cron job in place with the requested schedule and skill.
HA-15hermesagent-20020Failed to prepare the pinned Hermes runtime inside the verifier.
HA-16hermesagent-203020Hermes failed to send the message to the correct named target.
HA-17hermesagent-207020Hermes failed the parallel delegation scenario.
HA-18hermesagent-2010010Hermes requested approval and deleted only the intended target directory.
HA-19hermesagent-207020Hermes retried deployment partially, but the corrective-action trace or final success was incomplete.
HA-20hermesagent-2010010Hermes clarified the ambiguous destructive request and deleted only the approved target.
数据来源与校验Provenance and verification
  • toolcall-15 · toolcall-15-2026-07-20T07-34-55.537Z-ab7abea2 · SHA-256 5e9ae309dc55a5a7…
  • bugfind-15 · bugfind-15-2026-07-20T07-40-31.420Z-47089227 · SHA-256 fbd481f99aa9b583…
  • hermesagent-20 · hermesagent-20-2026-07-20T08-01-46.715Z-c037d50d · SHA-256 016171cefaf8d7c9…