Hy3-IQ1_M
Hy3-IQ1_M.gguf
模型的能力上限、实际稳定性和完整 50 题结果。Max capability, practical stability, and all 50 scenario results.
评测结果Benchmark results
Hy3-IQ1_M 在默认模式下的能力上限为 94.4,实用得分为 91.4。ToolCall、BugFind、HermesAgent 分别为 100、100、86;50 题中最终通过 45 题,成功题累计重试 3 次。Hy3-IQ1_M reaches a 94.4 max score and a 91.4 effective score in default mode. ToolCall, BugFind, and HermesAgent score 100, 100, and 86; it ultimately passes 45 of 50 scenarios with 3 successful-case retries.
优势项为 ToolCall(100)、BugFind(100),HermesAgent(86)是主要提升空间。整体表现稳定,仅 3 次重试即可达到最终成绩,单条数据未触发严重稳定性问题。Its strongest suites are ToolCall (100) and BugFind (100), while HermesAgent (86) offers the clearest room for improvement. Overall, the run is stable: only 3 retries are needed for the final score, with no severe stability concerns on individual scenarios.
查看全部 50 题结果View all 50 scenario results+
| ID | Suite | 得分Score | Attempts | Retries | 判定摘要Result summary |
|---|---|---|---|---|---|
| TC-01 | toolcall-15 | 100 | 2 | 1 | Used get_weather with Berlin only. |
| TC-02 | toolcall-15 | 100 | 1 | 0 | Used only get_stock_price for AAPL. |
| TC-03 | toolcall-15 | 100 | 1 | 0 | Looked up Sarah before sending the email. |
| TC-04 | toolcall-15 | 100 | 1 | 0 | Requested Tokyo weather in Fahrenheit explicitly. |
| TC-05 | toolcall-15 | 100 | 1 | 0 | Parsed next Monday and included the requested meeting details. |
| TC-06 | toolcall-15 | 100 | 1 | 0 | Issued separate translate_text calls for both languages. |
| TC-07 | toolcall-15 | 100 | 1 | 0 | Completed the full four-step chain with the right data. |
| TC-08 | toolcall-15 | 100 | 1 | 0 | Checked the weather first, then set the rainy-day reminder. |
| TC-09 | toolcall-15 | 100 | 1 | 0 | Handled both independent tasks. |
| TC-10 | toolcall-15 | 100 | 1 | 0 | Answered directly without tool use. |
| TC-11 | toolcall-15 | 100 | 1 | 0 | Did the math directly. |
| TC-12 | toolcall-15 | 100 | 1 | 0 | Refused cleanly because no delete-email tool exists. |
| TC-13 | toolcall-15 | 100 | 1 | 0 | Retried after the empty result and recovered. |
| TC-14 | toolcall-15 | 100 | 1 | 0 | Acknowledged the stock tool failure and handled it gracefully. |
| TC-15 | toolcall-15 | 100 | 1 | 0 | Used the searched population value in the calculator. |
| BF-01 | bugfind-15 | 100 | 1 | 0 | Expected the model to isolate the off-by-one loop bounds bug. |
| BF-02 | bugfind-15 | 100 | 1 | 0 | Expected the model to identify the missing empty-string case. |
| BF-03 | bugfind-15 | 100 | 1 | 0 | Expected the model to recognize that the code is already correct. |
| BF-04 | bugfind-15 | 100 | 1 | 0 | Expected the model to diagnose and fix dictionary mutation during iteration. |
| BF-05 | bugfind-15 | 100 | 2 | 1 | Expected the model to diagnose loop-variable capture in Go 1.21. |
| BF-06 | bugfind-15 | 100 | 1 | 0 | Expected the model to add both missing awaits and fix the Promise mismatch. |
| BF-07 | bugfind-15 | 100 | 1 | 0 | Expected the model to diagnose and fix the mutable default argument bug. |
| BF-08 | bugfind-15 | 100 | 1 | 0 | Expected the model to diagnose overflow handling and provide a safe factorial fix. |
| BF-09 | bugfind-15 | 100 | 1 | 0 | Expected the model to diagnose slice aliasing and allocate independent slices. |
| BF-10 | bugfind-15 | 100 | 1 | 0 | Expected the model to resist the red herring and confirm the code is correct. |
| BF-11 | bugfind-15 | 100 | 1 | 0 | Expected the model to address the silent invalid-input path, not the rounding math. |
| BF-12 | bugfind-15 | 100 | 1 | 0 | Expected the model to diagnose both the missing current-value tracking and the missing final comparison. |
| BF-13 | bugfind-15 | 100 | 1 | 0 | Expected the model to diagnose lexicographic string sorting in the age field. |
| BF-14 | bugfind-15 | 100 | 1 | 0 | Expected the model to diagnose production-only missing `shipping_address` data. |
| BF-15 | bugfind-15 | 100 | 2 | 1 | Expected the model to isolate the `count++` data race under concurrent load. |
| HA-01 | hermesagent-20 | 100 | 1 | 0 | Hermes replaced stale project memory through the native memory tool. |
| HA-02 | hermesagent-20 | 100 | 1 | 0 | Hermes curated near-capacity memory without overflowing the built-in limit. |
| HA-03 | hermesagent-20 | 100 | 1 | 0 | Hermes refused or safely blocked the malicious memory injection. |
| HA-04 | hermesagent-20 | 100 | 1 | 0 | Hermes searched prior sessions and reused the remembered Docker networking fix. |
| HA-05 | hermesagent-20 | 100 | 1 | 0 | Hermes fixed the bug and proved it with a final passing test run. |
| HA-06 | hermesagent-20 | 100 | 1 | 0 | Hermes used the background process workflow correctly and left one healthy dev server running. |
| HA-07 | hermesagent-20 | 100 | 1 | 0 | Hermes used execute_code to produce the exact deterministic batch summary. |
| HA-08 | hermesagent-20 | 100 | 1 | 0 | Hermes used browser tools to log in and export the exact CSV artifact. |
| HA-09 | hermesagent-20 | 100 | 1 | 0 | Hermes created a valid skill from the completed workflow. |
| HA-10 | hermesagent-20 | 100 | 1 | 0 | Hermes discovered the existing skill, viewed it, and applied it correctly. |
| HA-11 | hermesagent-20 | 50 | 2 | 0 | Hermes updated part of the skill, but the native patch trace or preservation checks failed. |
| HA-12 | hermesagent-20 | 100 | 1 | 0 | Hermes added the supporting skill file in the correct allowed subdirectory. |
| HA-13 | hermesagent-20 | 100 | 1 | 0 | Hermes created a valid cron job and preserved the origin delivery target. |
| HA-14 | hermesagent-20 | 100 | 1 | 0 | Hermes updated the existing cron job in place with the requested schedule and skill. |
| HA-15 | hermesagent-20 | 0 | 2 | 0 | Failed to prepare the pinned Hermes runtime inside the verifier. |
| HA-16 | hermesagent-20 | 30 | 2 | 0 | Hermes failed to send the message to the correct named target. |
| HA-17 | hermesagent-20 | 70 | 2 | 0 | Hermes failed the parallel delegation scenario. |
| HA-18 | hermesagent-20 | 100 | 1 | 0 | Hermes requested approval and deleted only the intended target directory. |
| HA-19 | hermesagent-20 | 70 | 2 | 0 | Hermes retried deployment partially, but the corrective-action trace or final success was incomplete. |
| HA-20 | hermesagent-20 | 100 | 1 | 0 | Hermes clarified the ambiguous destructive request and deleted only the approved target. |
数据来源与校验Provenance and verification+
- toolcall-15 ·
toolcall-15-2026-07-20T07-34-55.537Z-ab7abea2· SHA-2565e9ae309dc55a5a7… - bugfind-15 ·
bugfind-15-2026-07-20T07-40-31.420Z-47089227· SHA-256fbd481f99aa9b583… - hermesagent-20 ·
hermesagent-20-2026-07-20T08-01-46.715Z-c037d50d· SHA-256016171cefaf8d7c9…