⚡ August 05, 2026¶
Generated: 2026-08-05 12:14 UTC
Total Duration: 1h 28m 4s
Iterations: 3
Judge (classifier) model: gpt-4.1
Fast Benchmark
Markers: regression or benchmark
Schedule: Weekly (Sunday 2 AM UTC)
Purpose: Quick regression tests to catch breaking changes
HolmesGPT is continuously evaluated against real-world Kubernetes and cloud troubleshooting scenarios.
If you find scenarios that HolmesGPT does not perform well on, please consider adding them as evals to the benchmark.
Model Accuracy Comparison¶
| Model | Pass | Fail | Skip/Error | Total | Success Rate |
|---|---|---|---|---|---|
| gpt-5.6-luna | 45 | 18 | 0 | 63 | 🟡 71% (45/63) |
| gpt-5.6-sol | 54 | 9 | 0 | 63 | 🟡 86% (54/63) |
| gpt-5.6-terra | 52 | 11 | 0 | 63 | 🟡 83% (52/63) |
| opus-4.6 | 57 | 6 | 0 | 63 | 🟡 90% (57/63) |
| opus-4.8 | 55 | 8 | 0 | 63 | 🟡 87% (55/63) |
| opus-5 | 59 | 4 | 0 | 63 | 🟡 94% (59/63) |
| sonnet-5 | 52 | 11 | 0 | 63 | 🟡 83% (52/63) |
Model Cost Comparison¶
| Model | Tests | Avg Cost | Min Cost | Max Cost | Total Cost |
|---|---|---|---|---|---|
| gpt-5.6-luna | 63 | $0.01 | $0.00 | $0.03 | $0.53 |
| gpt-5.6-sol | 63 | $0.26 | $0.01 | $1.38 | $16.65 |
| gpt-5.6-terra | 63 | $0.07 | $0.00 | $0.24 | $4.54 |
| opus-4.6 | 63 | $0.27 | $0.08 | $2.15 | $16.73 |
| opus-4.8 | 63 | $0.27 | $0.10 | $1.26 | $17.15 |
| opus-5 | 63 | $0.39 | $0.10 | $3.90 | $24.54 |
| sonnet-5 | 63 | $0.12 | $0.04 | $0.53 | $7.52 |
Model Latency Comparison¶
| Model | Avg (s) | Min (s) | Max (s) | P50 (s) | P95 (s) |
|---|---|---|---|---|---|
| gpt-5.6-luna | 28.8 | 4.4 | 63.0 | 27.3 | 52.9 |
| gpt-5.6-sol | 44.2 | 5.1 | 171.9 | 42.2 | 91.2 |
| gpt-5.6-terra | 25.5 | 4.0 | 71.8 | 22.4 | 58.1 |
| opus-4.6 | 57.1 | 5.9 | 482.0 | 33.4 | 102.9 |
| opus-4.8 | 44.3 | 4.6 | 356.2 | 30.4 | 109.5 |
| opus-5 | 66.9 | 4.7 | 883.2 | 44.5 | 110.1 |
| sonnet-5 | 42.7 | 4.6 | 197.6 | 28.9 | 100.1 |
⚠️ Note: 1 test(s) excluded from latency calculations due to throttling/timeout errors (sonnet-5: 1)
Performance by Tag¶
Success rate by test category and model:
| Tag | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | opus-4.6 | opus-4.8 | opus-5 | sonnet-5 | Warnings |
|---|---|---|---|---|---|---|---|---|
| benchmark | 🟡 72% (13/18) | 🟡 78% (14/18) | 🟡 72% (13/18) | 🟡 83% (15/18) | 🟡 78% (14/18) | 🟡 94% (17/18) | 🟡 72% (13/18) | |
| context_window | 🟢 100% (6/6) | 🟡 83% (⅚) | 🟡 83% (⅚) | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟡 83% (⅚) | |
| counting | 🟡 33% (2/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟡 83% (⅚) | 🟢 100% (6/6) | |
| datetime | 🟢 100% (9/9) | 🟡 89% (8/9) | 🟡 67% (6/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟡 89% (8/9) | |
| easy | 🟡 54% (13/24) | 🟡 79% (19/24) | 🟡 75% (18/24) | 🟡 88% (21/24) | 🟡 83% (20/24) | 🟡 92% (22/24) | 🟡 75% (18/24) | |
| elasticsearch | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | |
| grafana | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | |
| hard | 🟡 50% (3/6) | 🟡 67% (4/6) | 🟡 50% (3/6) | 🟡 50% (3/6) | 🟡 50% (3/6) | 🟢 100% (6/6) | 🟡 50% (3/6) | |
| kubernetes | 🟡 67% (20/30) | 🟡 80% (24/30) | 🟡 83% (25/30) | 🟡 90% (27/30) | 🟡 83% (25/30) | 🟡 87% (26/30) | 🟡 77% (23/30) | |
| logs | 🟡 56% (10/18) | 🟡 67% (12/18) | 🟡 61% (11/18) | 🟡 72% (13/18) | 🟡 67% (12/18) | 🟡 83% (15/18) | 🟡 61% (11/18) | |
| loki | 🟡 50% (3/6) | 🟡 50% (3/6) | 🟡 50% (3/6) | 🟡 67% (4/6) | 🟡 50% (3/6) | 🟡 50% (3/6) | 🟡 50% (3/6) | |
| medium | 🟡 93% (25/27) | 🟡 93% (25/27) | 🟡 93% (25/27) | 🟢 100% (27/27) | 🟡 96% (26/27) | 🟡 96% (26/27) | 🟡 93% (25/27) | |
| metrics | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | |
| multi-cluster | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | |
| network | 🔴 0% (0/3) | 🟡 33% (⅓) | 🟢 100% (3/3) | 🟡 67% (⅔) | 🟡 33% (⅓) | 🟢 100% (3/3) | 🔴 0% (0/3) | |
| one-test | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | 🟢 100% (3/3) | |
| port-forward | 🟡 67% (6/9) | 🟡 67% (6/9) | 🟡 67% (6/9) | 🟡 78% (7/9) | 🟡 67% (6/9) | 🟡 67% (6/9) | 🟡 67% (6/9) | |
| question-answer | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | 🟢 100% (6/6) | |
| regression | 🟡 71% (32/45) | 🟡 89% (40/45) | 🟡 87% (39/45) | 🟡 93% (42/45) | 🟡 91% (41/45) | 🟡 93% (42/45) | 🟡 87% (39/45) | |
| skills | 🟡 61% (22/36) | 🟡 86% (31/36) | 🟡 83% (30/36) | 🟡 92% (33/36) | 🟡 89% (32/36) | 🟡 92% (33/36) | 🟡 83% (30/36) | |
| transparency | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | 🟢 100% (9/9) | |
| Overall | 🟡 71% (45/63) | 🟡 86% (54/63) | 🟡 83% (52/63) | 🟡 90% (57/63) | 🟡 87% (55/63) | 🟡 94% (59/63) | 🟡 83% (52/63) |
Raw Results¶
Status of all evaluations across models. Color coding:
- 🟢 Passing 100% (stable)
- 🟡 Passing 1-99%
- 🔴 Passing 0% (failing)
- 🔧 Mock data failure (missing or invalid test data)
- ⚠️ Setup failure (environment/infrastructure issue)
- ⏱️ Timeout or rate limit error
- ⏭️ Test skipped (e.g., known issue or precondition not met)
Detailed Raw Results¶
| Eval ID | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | opus-4.6 | opus-4.8 | opus-5 | sonnet-5 |
|---|---|---|---|---|---|---|---|
| 09_crashpod 🔗 | 🟢 100% (3/3) / ⏱️ 16.6s / 💰 $0.00 | 🟢 100% (3/3) / ⏱️ 41.9s / 💰 $0.22 | 🟢 100% (3/3) / ⏱️ 16.2s / 💰 $0.04 | 🟢 100% (3/3) / ⏱️ 30.3s / 💰 $0.17 | 🟢 100% (3/3) / ⏱️ 22.5s / 💰 $0.17 | 🟢 100% (3/3) / ⏱️ 31.0s / 💰 $0.22 | 🟢 100% (3/3) / ⏱️ 22.8s / 💰 $0.08 |
| 100a_loki_historical_logs 🔗 | 🟡 67% (⅔) / ⏱️ 41.8s / 💰 $0.01 | 🟡 67% (⅔) / ⏱️ 72.0s / 💰 $0.47 | 🟡 67% (⅔) / ⏱️ 31.5s / 💰 $0.10 | 🟢 100% (3/3) / ⏱️ 108.4s / 💰 $0.40 | 🟡 67% (⅔) / ⏱️ 73.6s / 💰 $0.44 | 🟡 67% (⅔) / ⏱️ 127.3s / 💰 $0.70 | 🟡 67% (⅔) / ⏱️ 71.4s / 💰 $0.19 |
| 101_loki_historical_logs_pod_deleted 🔗 | 🟡 33% (⅓) / ⏱️ 44.0s / 💰 $0.01 | 🟡 33% (⅓) / ⏱️ 86.5s / 💰 $0.39 | 🟡 33% (⅓) / ⏱️ 34.4s / 💰 $0.08 | 🟡 33% (⅓) / ⏱️ 310.3s / 💰 $1.39 | 🟡 33% (⅓) / ⏱️ 172.8s / 💰 $0.67 | 🟡 33% (⅓) / ⏱️ 384.1s / 💰 $1.79 | 🟡 33% (⅓) / ⏱️ 85.3s / 💰 $0.18 |
| 108_logs_nearby_lines 🔗 | 🔴 0% (0/3) / ⏱️ 36.2s / 💰 $0.01 | 🟡 33% (⅓) / ⏱️ 48.8s / 💰 $0.29 | 🔴 0% (0/3) / ⏱️ 29.8s / 💰 $0.08 | 🔴 0% (0/3) / ⏱️ 49.4s / 💰 $0.26 | 🔴 0% (0/3) / ⏱️ 62.1s / 💰 $0.36 | 🟢 100% (3/3) / ⏱️ 65.9s / 💰 $0.49 | 🔴 0% (0/3) / ⏱️ 63.3s / 💰 $0.11 |
| 112_find_pvcs_by_uuid 🔗 | 🟢 100% (3/3) / ⏱️ 12.6s / 💰 $0.00 | 🟢 100% (3/3) / ⏱️ 20.6s / 💰 $0.10 | 🟢 100% (3/3) / ⏱️ 9.2s / 💰 $0.02 | 🟢 100% (3/3) / ⏱️ 25.5s / 💰 $0.14 | 🟢 100% (3/3) / ⏱️ 20.0s / 💰 $0.17 | 🟢 100% (3/3) / ⏱️ 30.8s / 💰 $0.24 | 🟢 100% (3/3) / ⏱️ 13.0s / 💰 $0.06 |
| 12_job_crashing 🔗 | 🟢 100% (3/3) / ⏱️ 18.0s / 💰 $0.00 | 🟢 100% (3/3) / ⏱️ 21.3s / 💰 $0.11 | 🟢 100% (3/3) / ⏱️ 16.4s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 24.8s / 💰 $0.14 | 🟢 100% (3/3) / ⏱️ 29.4s / 💰 $0.21 | 🟢 100% (3/3) / ⏱️ 26.0s / 💰 $0.20 | 🟢 100% (3/3) / ⏱️ 19.9s / 💰 $0.07 |
| 176_network_policy_blocking_traffic_no_skills 🔗 | 🔴 0% (0/3) / ⏱️ 51.1s / 💰 $0.02 | 🟡 33% (⅓) / ⏱️ 57.4s / 💰 $0.62 | 🟢 100% (3/3) / ⏱️ 22.4s / 💰 $0.06 | 🟡 67% (⅔) / ⏱️ 40.6s / 💰 $0.21 | 🟡 33% (⅓) / ⏱️ 89.4s / 💰 $0.49 | 🟢 100% (3/3) / ⏱️ 50.4s / 💰 $0.28 | 🔴 0% (0/3) / ⏱️ 150.0s / 💰 $0.40 |
| 179_grafana_big_dashboard_query 🔗 | 🟢 100% (3/3) / ⏱️ 17.8s / 💰 $0.01 | 🟢 100% (3/3) / ⏱️ 15.2s / 💰 $0.10 | 🟢 100% (3/3) / ⏱️ 9.2s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 18.5s / 💰 $0.15 | 🟢 100% (3/3) / ⏱️ 14.1s / 💰 $0.22 | 🟢 100% (3/3) / ⏱️ 31.8s / 💰 $0.25 | 🟢 100% (3/3) / ⏱️ 15.7s / 💰 $0.08 |
| 227_count_configmaps_per_namespace[0] 🔗 | 🟡 33% (⅓) / ⏱️ 24.1s / 💰 $0.01 | 🟢 100% (3/3) / ⏱️ 16.9s / 💰 $0.08 | 🟢 100% (3/3) / ⏱️ 13.8s / 💰 $0.04 | 🟢 100% (3/3) / ⏱️ 18.0s / 💰 $0.12 | 🟢 100% (3/3) / ⏱️ 21.9s / 💰 $0.22 | 🟡 67% (⅔) / ⏱️ 31.0s / 💰 $0.24 | 🟢 100% (3/3) / ⏱️ 21.3s / 💰 $0.12 |
| 243_pod_names_contain_service 🔗 | 🟢 100% (3/3) / ⏱️ 19.7s / 💰 $0.00 | 🟢 100% (3/3) / ⏱️ 30.9s / 💰 $0.14 | 🟢 100% (3/3) / ⏱️ 21.0s / 💰 $0.04 | 🟢 100% (3/3) / ⏱️ 31.7s / 💰 $0.15 | 🟢 100% (3/3) / ⏱️ 22.3s / 💰 $0.17 | 🟢 100% (3/3) / ⏱️ 35.1s / 💰 $0.17 | 🟢 100% (3/3) / ⏱️ 24.5s / 💰 $0.07 |
| 24_misconfigured_pvc 🔗 | 🟡 33% (⅓) / ⏱️ 23.2s / 💰 $0.01 | 🟡 67% (⅔) / ⏱️ 45.5s / 💰 $0.27 | 🟡 33% (⅓) / ⏱️ 22.0s / 💰 $0.06 | 🟢 100% (3/3) / ⏱️ 31.6s / 💰 $0.18 | 🟢 100% (3/3) / ⏱️ 26.5s / 💰 $0.19 | 🟢 100% (3/3) / ⏱️ 33.7s / 💰 $0.23 | 🟡 67% (⅔) / ⏱️ 27.9s / 💰 $0.08 |
| 254_elasticsearch_dr_test_log_check 🔗 | 🟢 100% (3/3) / ⏱️ 46.8s / 💰 $0.01 | 🟢 100% (3/3) / ⏱️ 78.4s / 💰 $0.43 | 🟢 100% (3/3) / ⏱️ 56.5s / 💰 $0.16 | 🟢 100% (3/3) / ⏱️ 84.8s / 💰 $0.34 | 🟢 100% (3/3) / ⏱️ 51.2s / 💰 $0.26 | 🟢 100% (3/3) / ⏱️ 83.0s / 💰 $0.37 | 🟢 100% (3/3) / ⏱️ 58.7s / 💰 $0.13 |
| 259_wrong_cluster_logs_confusion 🔗 | 🟢 100% (3/3) / ⏱️ 52.1s / 💰 $0.01 | 🟢 100% (3/3) / ⏱️ 89.9s / 💰 $0.55 | 🟢 100% (3/3) / ⏱️ 51.4s / 💰 $0.14 | 🟢 100% (3/3) / ⏱️ 84.7s / 💰 $0.32 | 🟢 100% (3/3) / ⏱️ 62.0s / 💰 $0.31 | 🟢 100% (3/3) / ⏱️ 77.6s / 💰 $0.42 | 🟢 100% (3/3) / ⏱️ 65.9s / 💰 $0.15 |
| 260_global_es_remote_cluster_logs 🔗 | 🟢 100% (3/3) / ⏱️ 44.2s / 💰 $0.01 | 🟢 100% (3/3) / ⏱️ 48.3s / 💰 $0.20 | 🟢 100% (3/3) / ⏱️ 58.5s / 💰 $0.17 | 🟢 100% (3/3) / ⏱️ 90.2s / 💰 $0.33 | 🟢 100% (3/3) / ⏱️ 52.0s / 💰 $0.25 | 🟢 100% (3/3) / ⏱️ 72.3s / 💰 $0.32 | 🟢 100% (3/3) / ⏱️ 56.3s / 💰 $0.11 |
| 283_todowrite_multistep_audit 🔗 | 🟢 100% (3/3) / ⏱️ 26.9s / 💰 $0.01 | 🟢 100% (3/3) / ⏱️ 29.8s / 💰 $0.22 | 🟢 100% (3/3) / ⏱️ 24.3s / 💰 $0.08 | 🟢 100% (3/3) / ⏱️ 27.2s / 💰 $0.16 | 🟢 100% (3/3) / ⏱️ 21.1s / 💰 $0.18 | 🟢 100% (3/3) / ⏱️ 27.1s / 💰 $0.20 | 🟢 100% (3/3) / ⏱️ 25.4s / 💰 $0.08 |
| 43_current_datetime_from_prompt 🔗 | 🟢 100% (3/3) / ⏱️ 4.8s / 💰 $0.00 | 🟢 100% (3/3) / ⏱️ 5.9s / 💰 $0.01 | 🟡 33% (⅓) / ⏱️ 5.5s / 💰 $0.00 | 🟢 100% (3/3) / ⏱️ 6.9s / 💰 $0.08 | 🟢 100% (3/3) / ⏱️ 5.5s / 💰 $0.10 | 🟢 100% (3/3) / ⏱️ 4.8s / 💰 $0.10 | 🟢 100% (3/3) / ⏱️ 5.1s / 💰 $0.04 |
| 51_logs_summarize_errors 🔗 | 🟡 33% (⅓) / ⏱️ 17.6s / 💰 $0.00 | 🟢 100% (3/3) / ⏱️ 20.8s / 💰 $0.11 | 🟢 100% (3/3) / ⏱️ 11.4s / 💰 $0.02 | 🟢 100% (3/3) / ⏱️ 23.0s / 💰 $0.13 | 🟢 100% (3/3) / ⏱️ 22.4s / 💰 $0.18 | 🟢 100% (3/3) / ⏱️ 27.9s / 💰 $0.23 | 🟢 100% (3/3) / ⏱️ 20.2s / 💰 $0.07 |
| 61_exact_match_counting 🔗 | 🟡 33% (⅓) / ⏱️ 6.7s / 💰 $0.00 | 🟢 100% (3/3) / ⏱️ 6.4s / 💰 $0.03 | 🟢 100% (3/3) / ⏱️ 5.6s / 💰 $0.02 | 🟢 100% (3/3) / ⏱️ 9.7s / 💰 $0.09 | 🟢 100% (3/3) / ⏱️ 7.8s / 💰 $0.20 | 🟢 100% (3/3) / ⏱️ 12.0s / 💰 $0.12 | 🟢 100% (3/3) / ⏱️ 8.7s / 💰 $0.05 |
| 73a_time_window_anomaly 🔗 | 🟢 100% (3/3) / ⏱️ 32.9s / 💰 $0.01 | 🟡 67% (⅔) / ⏱️ 62.2s / 💰 $0.53 | 🟡 67% (⅔) / ⏱️ 28.5s / 💰 $0.07 | 🟢 100% (3/3) / ⏱️ 59.3s / 💰 $0.25 | 🟢 100% (3/3) / ⏱️ 33.6s / 💰 $0.22 | 🟢 100% (3/3) / ⏱️ 94.6s / 💰 $0.64 | 🟡 67% (⅔) / ⏱️ 52.9s / 💰 $0.13 |
| 73b_time_window_anomaly 🔗 | 🟢 100% (3/3) / ⏱️ 26.5s / 💰 $0.01 | 🟢 100% (3/3) / ⏱️ 70.2s / 💰 $0.34 | 🟢 100% (3/3) / ⏱️ 31.0s / 💰 $0.08 | 🟢 100% (3/3) / ⏱️ 57.5s / 💰 $0.27 | 🟢 100% (3/3) / ⏱️ 33.4s / 💰 $0.22 | 🟢 100% (3/3) / ⏱️ 92.9s / 💰 $0.59 | 🟢 100% (3/3) / ⏱️ 41.6s / 💰 $0.12 |
| 96_no_matching_skill 🔗 | 🟡 67% (⅔) / ⏱️ 42.0s / 💰 $0.02 | 🟢 100% (3/3) / ⏱️ 58.7s / 💰 $0.36 | 🟢 100% (3/3) / ⏱️ 36.3s / 💰 $0.18 | 🟢 100% (3/3) / ⏱️ 66.6s / 💰 $0.32 | 🟢 100% (3/3) / ⏱️ 85.8s / 💰 $0.49 | 🟢 100% (3/3) / ⏱️ 65.0s / 💰 $0.37 | 🟢 100% (3/3) / ⏱️ 70.9s / 💰 $0.18 |
Results are automatically generated and updated weekly. View full traces and detailed analysis in Braintrust experiment: ci-benchmark-30998420319.