5.2 Evaluation Results and Leaderboards¶
This generated report is the complete static comparison of the committed evaluation and judgment snapshots. It complements the evaluation methodology, the narrative comparison, the dataset complexity ladder, and the raw result snapshots.
Model provenance. Recorded judge models: gemma4:31b, qwen3.6:latest. Recorded Ragas evaluator models: mistral-small3.2:24b.
Active plugin generation role models: qwen3.8:latest. Active LightRAG role models: qwen3.8:latest. Snapshot model names remain unchanged because they identify the systems that produced the recorded answers and scores.
1. Reading the Results¶
The default overall order is the dataset-macro judge mean: each measured dataset contributes one equally weighted judge mean. Query-weighted judge mean is shown separately and weights each evaluated query equally. No composite score combines quality, coverage, latency, or operational reliability.
Higher judge, answer-relevancy, faithfulness, eligible-coverage, successful-response, and
per-query-win values are better. Lower ranks, disagreement, latency, error rate,
errors, and timeouts are better. Judge coverage is evaluated judge questions over all
judge questions. Ragas coverage is evaluated rows over eligible rows (total rows -
ineligible), while each metric's total rows, ineligible rows, evaluator errors, and
timeouts remain separate columns. N/A means no value was recorded or no rows were
eligible and carries an empty machine sort value. Faithfulness ineligible rows are not
failures and are never coerced to zero. Ragas evaluator errors and timeouts also remain
separate from response errors and timeouts.
Base approaches and flavor aliases are intentionally separate tiers. A flavor identifies its base family but cannot occupy a base-approach rank.
2. Overall Base-Approach Leaderboard¶
| Overall judge rank | Approach | Maturity | Dataset-macro judge | Query-weighted judge | Judge coverage | Judge errors | Judge gemma4:31b | Judge gemma4:31b coverage | Judge qwen3.6:latest | Judge qwen3.6:latest coverage | Judge disagreement | Judge disagreement comparisons | Mean dataset rank | Best dataset rank | Worst dataset rank | Per-query wins | Answer relevancy mean | Answer relevancy coverage (eligible) | Answer relevancy total rows | Answer relevancy ineligible | Answer relevancy errors | Answer relevancy timeouts | Faithfulness mean | Faithfulness coverage (eligible) | Faithfulness total rows | Faithfulness ineligible | Faithfulness errors | Faithfulness timeouts | Mean latency (ms) | Successful | Attempted | Error rate | Errors | Timeouts |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | contextual-rag | canonical | 3.757 | 3.800 | 20 / 20 (100.00%) | 0 | 3.550 | 20 / 20 (100.00%) | 4.050 | 20 / 20 (100.00%) | 0.600 | 20 | 2.000 | 1 | 3 | 4 | 0.872 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.417 | 12 / 20 (60.00%) | 20 | 0 | 8 | 0 | 17143.75 | 20 | 20 | 0.00% | 0 | 0 |
| 2 | lazy-graph-rag | experimental | 3.743 | 3.800 | 20 / 20 (100.00%) | 0 | 3.650 | 20 / 20 (100.00%) | 3.950 | 20 / 20 (100.00%) | 0.500 | 20 | 2.000 | 1 | 3 | 4 | 0.851 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.433 | 13 / 20 (65.00%) | 20 | 0 | 7 | 0 | 6065.35 | 20 | 20 | 0.00% | 0 | 0 |
| 2 | vanilla-rag | canonical | 3.743 | 3.775 | 20 / 20 (100.00%) | 0 | 3.650 | 20 / 20 (100.00%) | 3.900 | 20 / 20 (100.00%) | 0.850 | 20 | 2.000 | 1 | 3 | 4 | 0.798 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.399 | 14 / 20 (70.00%) | 20 | 0 | 6 | 0 | 5637.90 | 20 | 20 | 0.00% | 0 | 0 |
| 4 | hybrid-rag | canonical | 3.514 | 3.525 | 20 / 20 (100.00%) | 0 | 3.500 | 20 / 20 (100.00%) | 3.550 | 20 / 20 (100.00%) | 0.750 | 20 | 4.000 | 2 | 6 | 4 | 0.868 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.379 | 15 / 20 (75.00%) | 20 | 0 | 5 | 0 | 16044.75 | 20 | 20 | 0.00% | 0 | 0 |
| 5 | graph-rag | canonical | 2.931 | 2.900 | 20 / 20 (100.00%) | 0 | 2.600 | 20 / 20 (100.00%) | 3.200 | 20 / 20 (100.00%) | 0.900 | 20 | 5.667 | 5 | 7 | 4 | 0.805 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | N/A | N/A | 20 | 20 | 0 | 0 | 15134.00 | 20 | 20 | 0.00% | 0 | 0 |
| 6 | n8n-adaptive-rag | canonical | 2.924 | 2.875 | 20 / 20 (100.00%) | 0 | 2.750 | 20 / 20 (100.00%) | 3.000 | 20 / 20 (100.00%) | 0.850 | 20 | 4.667 | 2 | 6 | 0 | 0.779 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.388 | 9 / 20 (45.00%) | 20 | 0 | 11 | 0 | 6279.25 | 20 | 20 | 0.00% | 0 | 0 |
| 7 | agentic-rag | canonical | 2.701 | 2.675 | 20 / 20 (100.00%) | 0 | 2.550 | 20 / 20 (100.00%) | 2.800 | 20 / 20 (100.00%) | 0.850 | 20 | 5.000 | 2 | 7 | 0 | 0.729 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.443 | 9 / 20 (45.00%) | 20 | 0 | 11 | 0 | 27434.60 | 20 | 20 | 0.00% | 0 | 0 |
3. Base Approaches by Dataset¶
| Dataset | Complexity | Approach | Maturity | Judge rank | Judge mean | Judge coverage | Judge errors | Judge gemma4:31b | Judge gemma4:31b coverage | Judge qwen3.6:latest | Judge qwen3.6:latest coverage | Judge disagreement | Judge disagreement comparisons | Per-query wins | Answer relevancy rank | Answer relevancy mean | Answer relevancy coverage (eligible) | Answer relevancy total rows | Answer relevancy ineligible | Answer relevancy errors | Answer relevancy timeouts | Faithfulness rank | Faithfulness mean | Faithfulness coverage (eligible) | Faithfulness total rows | Faithfulness ineligible | Faithfulness errors | Faithfulness timeouts | Latency rank | Mean latency (ms) | Successful | Attempted | Error rate | Errors | Timeouts |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| baseline_curated | 1 | vanilla-rag | canonical | 1 | 4.167 | 6 / 6 (100.00%) | 0 | 4.333 | 6 / 6 (100.00%) | 4.000 | 6 / 6 (100.00%) | 1.000 | 6 | 0 | 6 | 0.708 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 1 | 0.672 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 2 | 3833.83 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | hybrid-rag | canonical | 2 | 4.000 | 6 / 6 (100.00%) | 0 | 4.333 | 6 / 6 (100.00%) | 3.667 | 6 / 6 (100.00%) | 1.000 | 6 | 1 | 2 | 0.869 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 3 | 0.645 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 6 | 17682.83 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | contextual-rag | canonical | 3 | 3.917 | 6 / 6 (100.00%) | 0 | 3.500 | 6 / 6 (100.00%) | 4.333 | 6 / 6 (100.00%) | 0.833 | 6 | 0 | 1 | 0.894 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 6 | 0.300 | 5 / 6 (83.33%) | 6 | 0 | 1 | 0 | 7 | 18016.33 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | lazy-graph-rag | experimental | 3 | 3.917 | 6 / 6 (100.00%) | 0 | 3.833 | 6 / 6 (100.00%) | 4.000 | 6 / 6 (100.00%) | 0.167 | 6 | 1 | 3 | 0.858 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 4 | 0.636 | 5 / 6 (83.33%) | 6 | 0 | 1 | 0 | 3 | 5514.00 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | graph-rag | canonical | 5 | 3.750 | 6 / 6 (100.00%) | 0 | 3.667 | 6 / 6 (100.00%) | 3.833 | 6 / 6 (100.00%) | 0.500 | 6 | 4 | 4 | 0.842 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | N/A | N/A | N/A | 6 | 6 | 0 | 0 | 5 | 12611.83 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | n8n-adaptive-rag | canonical | 6 | 3.333 | 6 / 6 (100.00%) | 0 | 3.667 | 6 / 6 (100.00%) | 3.000 | 6 / 6 (100.00%) | 0.667 | 6 | 0 | 5 | 0.733 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 5 | 0.564 | 5 / 6 (83.33%) | 6 | 0 | 1 | 0 | 1 | 2798.33 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | agentic-rag | canonical | 7 | 2.667 | 6 / 6 (100.00%) | 0 | 3.000 | 6 / 6 (100.00%) | 2.333 | 6 / 6 (100.00%) | 0.667 | 6 | 0 | 7 | 0.567 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 2 | 0.664 | 5 / 6 (83.33%) | 6 | 0 | 1 | 0 | 4 | 10773.67 | 6 | 6 | 0.00% | 0 | 0 |
| graph_native | 2 | lazy-graph-rag | experimental | 1 | 4.312 | 8 / 8 (100.00%) | 0 | 4.125 | 8 / 8 (100.00%) | 4.500 | 8 / 8 (100.00%) | 0.625 | 8 | 2 | 4 | 0.829 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 2 | 0.416 | 5 / 8 (62.50%) | 8 | 0 | 3 | 0 | 2 | 4935.75 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | contextual-rag | canonical | 2 | 4.188 | 8 / 8 (100.00%) | 0 | 4.000 | 8 / 8 (100.00%) | 4.375 | 8 / 8 (100.00%) | 0.625 | 8 | 2 | 3 | 0.841 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 1 | 0.607 | 5 / 8 (62.50%) | 8 | 0 | 3 | 0 | 5 | 11196.88 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | vanilla-rag | canonical | 3 | 4.062 | 8 / 8 (100.00%) | 0 | 3.750 | 8 / 8 (100.00%) | 4.375 | 8 / 8 (100.00%) | 0.875 | 8 | 3 | 2 | 0.851 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 4 | 0.279 | 5 / 8 (62.50%) | 8 | 0 | 3 | 0 | 1 | 4539.25 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | hybrid-rag | canonical | 4 | 3.625 | 8 / 8 (100.00%) | 0 | 3.375 | 8 / 8 (100.00%) | 3.875 | 8 / 8 (100.00%) | 0.750 | 8 | 1 | 1 | 0.859 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 3 | 0.303 | 6 / 8 (75.00%) | 8 | 0 | 2 | 0 | 4 | 9808.88 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | graph-rag | canonical | 5 | 2.625 | 8 / 8 (100.00%) | 0 | 2.125 | 8 / 8 (100.00%) | 3.125 | 8 / 8 (100.00%) | 1.250 | 8 | 0 | 5 | 0.796 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | N/A | N/A | N/A | 8 | 8 | 0 | 0 | 6 | 12473.88 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | agentic-rag | canonical | 6 | 2.438 | 8 / 8 (100.00%) | 0 | 2.125 | 8 / 8 (100.00%) | 2.750 | 8 / 8 (100.00%) | 0.625 | 8 | 0 | 6 | 0.740 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 5 | 0.000 | 1 / 8 (12.50%) | 8 | 0 | 7 | 0 | 7 | 29112.75 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | n8n-adaptive-rag | canonical | 6 | 2.438 | 8 / 8 (100.00%) | 0 | 2.125 | 8 / 8 (100.00%) | 2.750 | 8 / 8 (100.00%) | 0.625 | 8 | 0 | 6 | 0.740 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 5 | 0.000 | 1 / 8 (12.50%) | 8 | 0 | 7 | 0 | 3 | 5299.00 | 8 | 8 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | contextual-rag | canonical | 1 | 3.167 | 6 / 6 (100.00%) | 0 | 3.000 | 6 / 6 (100.00%) | 3.333 | 6 / 6 (100.00%) | 0.333 | 6 | 2 | 1 | 0.891 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 1 | 0.232 | 2 / 6 (33.33%) | 6 | 0 | 4 | 0 | 6 | 24200.33 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | agentic-rag | canonical | 2 | 3.000 | 6 / 6 (100.00%) | 0 | 2.667 | 6 / 6 (100.00%) | 3.333 | 6 / 6 (100.00%) | 1.333 | 6 | 0 | 3 | 0.877 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 2 | 0.224 | 3 / 6 (50.00%) | 6 | 0 | 3 | 0 | 7 | 41858.00 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | lazy-graph-rag | experimental | 2 | 3.000 | 6 / 6 (100.00%) | 0 | 2.833 | 6 / 6 (100.00%) | 3.167 | 6 / 6 (100.00%) | 0.667 | 6 | 1 | 5 | 0.873 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 4 | 0.121 | 3 / 6 (50.00%) | 6 | 0 | 3 | 0 | 1 | 8122.83 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | n8n-adaptive-rag | canonical | 2 | 3.000 | 6 / 6 (100.00%) | 0 | 2.667 | 6 / 6 (100.00%) | 3.333 | 6 / 6 (100.00%) | 1.333 | 6 | 0 | 3 | 0.877 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 2 | 0.224 | 3 / 6 (50.00%) | 6 | 0 | 3 | 0 | 3 | 11067.17 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | vanilla-rag | canonical | 2 | 3.000 | 6 / 6 (100.00%) | 0 | 2.833 | 6 / 6 (100.00%) | 3.167 | 6 / 6 (100.00%) | 0.667 | 6 | 1 | 6 | 0.819 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 5 | 0.052 | 3 / 6 (50.00%) | 6 | 0 | 3 | 0 | 2 | 8906.83 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | hybrid-rag | canonical | 6 | 2.917 | 6 / 6 (100.00%) | 0 | 2.833 | 6 / 6 (100.00%) | 3.000 | 6 / 6 (100.00%) | 0.500 | 6 | 2 | 2 | 0.878 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 6 | 0.000 | 3 / 6 (50.00%) | 6 | 0 | 3 | 0 | 5 | 22721.17 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | graph-rag | canonical | 7 | 2.417 | 6 / 6 (100.00%) | 0 | 2.167 | 6 / 6 (100.00%) | 2.667 | 6 / 6 (100.00%) | 0.833 | 6 | 0 | 7 | 0.780 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | N/A | N/A | N/A | 6 | 6 | 0 | 0 | 4 | 21203.00 | 6 | 6 | 0.00% | 0 | 0 |
4. Overall Flavor-Alias Leaderboard¶
| Overall judge rank | Flavor | Base family | Maturity | Dataset-macro judge | Query-weighted judge | Judge coverage | Judge errors | Judge gemma4:31b | Judge gemma4:31b coverage | Judge qwen3.6:latest | Judge qwen3.6:latest coverage | Judge disagreement | Judge disagreement comparisons | Mean dataset rank | Best dataset rank | Worst dataset rank | Per-query wins | Answer relevancy mean | Answer relevancy coverage (eligible) | Answer relevancy total rows | Answer relevancy ineligible | Answer relevancy errors | Answer relevancy timeouts | Faithfulness mean | Faithfulness coverage (eligible) | Faithfulness total rows | Faithfulness ineligible | Faithfulness errors | Faithfulness timeouts | Mean latency (ms) | Successful | Attempted | Error rate | Errors | Timeouts |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | hybrid-rag-fast | hybrid-rag | canonical | 3.792 | 3.775 | 20 / 20 (100.00%) | 0 | 3.600 | 20 / 20 (100.00%) | 3.950 | 20 / 20 (100.00%) | 0.650 | 20 | 2.667 | 1 | 5 | 2 | 0.797 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.389 | 16 / 20 (80.00%) | 20 | 0 | 4 | 0 | 4815.90 | 20 | 20 | 0.00% | 0 | 0 |
| 2 | hybrid-rag-high-recall | hybrid-rag | canonical | 3.757 | 3.800 | 20 / 20 (100.00%) | 0 | 3.500 | 20 / 20 (100.00%) | 4.100 | 20 / 20 (100.00%) | 0.600 | 20 | 3.667 | 1 | 7 | 2 | 0.877 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.420 | 12 / 20 (60.00%) | 20 | 0 | 8 | 0 | 23866.35 | 20 | 20 | 0.00% | 0 | 0 |
| 3 | lazy-graph-rag-wide | lazy-graph-rag | experimental | 3.729 | 3.725 | 20 / 20 (100.00%) | 0 | 3.650 | 20 / 20 (100.00%) | 3.800 | 20 / 20 (100.00%) | 0.450 | 20 | 3.667 | 1 | 7 | 2 | 0.818 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.468 | 11 / 20 (55.00%) | 20 | 0 | 9 | 0 | 7197.30 | 20 | 20 | 0.00% | 0 | 0 |
| 4 | vanilla-rag-wide | vanilla-rag | canonical | 3.694 | 3.725 | 20 / 20 (100.00%) | 0 | 3.400 | 20 / 20 (100.00%) | 4.050 | 20 / 20 (100.00%) | 0.750 | 20 | 3.000 | 2 | 5 | 4 | 0.839 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.455 | 15 / 20 (75.00%) | 20 | 0 | 5 | 0 | 8154.35 | 20 | 20 | 0.00% | 0 | 0 |
| 5 | contextual-rag-high-recall | contextual-rag | canonical | 3.597 | 3.575 | 20 / 20 (100.00%) | 0 | 3.400 | 20 / 20 (100.00%) | 3.750 | 20 / 20 (100.00%) | 0.450 | 20 | 4.333 | 2 | 7 | 0 | 0.855 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.378 | 12 / 20 (60.00%) | 20 | 0 | 8 | 0 | 24930.25 | 20 | 20 | 0.00% | 0 | 0 |
| 5 | lazy-graph-rag-fast | lazy-graph-rag | experimental | 3.597 | 3.600 | 20 / 20 (100.00%) | 0 | 3.300 | 20 / 20 (100.00%) | 3.900 | 20 / 20 (100.00%) | 0.700 | 20 | 4.333 | 3 | 5 | 3 | 0.846 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.424 | 14 / 20 (70.00%) | 20 | 0 | 6 | 0 | 4786.25 | 20 | 20 | 0.00% | 0 | 0 |
| 7 | lazy-graph-rag-balanced | lazy-graph-rag | experimental | 3.479 | 3.500 | 20 / 20 (100.00%) | 0 | 3.200 | 20 / 20 (100.00%) | 3.800 | 20 / 20 (100.00%) | 0.600 | 20 | 5.000 | 3 | 7 | 1 | 0.854 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.465 | 13 / 20 (65.00%) | 20 | 0 | 7 | 0 | 5752.70 | 20 | 20 | 0.00% | 0 | 0 |
| 8 | graph-rag-wide | graph-rag | canonical | 2.875 | 2.850 | 20 / 20 (100.00%) | 0 | 2.650 | 20 / 20 (100.00%) | 3.050 | 20 / 20 (100.00%) | 0.500 | 20 | 8.000 | 5 | 11 | 4 | 0.835 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | N/A | N/A | 20 | 20 | 0 | 0 | 18603.75 | 20 | 20 | 0.00% | 0 | 0 |
| 9 | graph-rag-rerank | graph-rag | canonical | 2.840 | 2.800 | 20 / 20 (100.00%) | 0 | 2.700 | 20 / 20 (100.00%) | 2.900 | 20 / 20 (100.00%) | 0.400 | 20 | 9.000 | 9 | 9 | 1 | 0.822 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | N/A | N/A | 20 | 20 | 0 | 0 | 21719.25 | 20 | 20 | 0.00% | 0 | 0 |
| 10 | n8n-adaptive-rag-default | n8n-adaptive-rag | canonical | 2.708 | 2.675 | 20 / 20 (100.00%) | 0 | 2.450 | 20 / 20 (100.00%) | 2.900 | 20 / 20 (100.00%) | 0.450 | 20 | 9.667 | 8 | 11 | 1 | 0.738 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.463 | 9 / 20 (45.00%) | 20 | 0 | 11 | 0 | 8158.95 | 20 | 20 | 0.00% | 0 | 0 |
| 11 | graph-rag-fast | graph-rag | canonical | 2.549 | 2.500 | 20 / 20 (100.00%) | 0 | 2.300 | 20 / 20 (100.00%) | 2.700 | 20 / 20 (100.00%) | 0.500 | 20 | 11.000 | 9 | 12 | 0 | 0.834 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | N/A | N/A | 20 | 20 | 0 | 0 | 12400.35 | 20 | 20 | 0.00% | 0 | 0 |
| 12 | agentic-rag-deeper | agentic-rag | canonical | 2.528 | 2.500 | 20 / 20 (100.00%) | 0 | 2.350 | 20 / 20 (100.00%) | 2.650 | 20 / 20 (100.00%) | 0.400 | 20 | 10.667 | 9 | 12 | 0 | 0.725 | 20 / 20 (100.00%) | 20 | 0 | 0 | 0 | 0.731 | 8 / 20 (40.00%) | 20 | 0 | 12 | 0 | 17397.75 | 20 | 20 | 0.00% | 0 | 0 |
5. Flavor Aliases by Dataset¶
| Dataset | Complexity | Flavor | Base family | Maturity | Judge rank | Judge mean | Judge coverage | Judge errors | Judge gemma4:31b | Judge gemma4:31b coverage | Judge qwen3.6:latest | Judge qwen3.6:latest coverage | Judge disagreement | Judge disagreement comparisons | Per-query wins | Answer relevancy rank | Answer relevancy mean | Answer relevancy coverage (eligible) | Answer relevancy total rows | Answer relevancy ineligible | Answer relevancy errors | Answer relevancy timeouts | Faithfulness rank | Faithfulness mean | Faithfulness coverage (eligible) | Faithfulness total rows | Faithfulness ineligible | Faithfulness errors | Faithfulness timeouts | Latency rank | Mean latency (ms) | Successful | Attempted | Error rate | Errors | Timeouts |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| baseline_curated | 1 | lazy-graph-rag-wide | lazy-graph-rag | experimental | 1 | 4.583 | 6 / 6 (100.00%) | 0 | 4.500 | 6 / 6 (100.00%) | 4.667 | 6 / 6 (100.00%) | 0.167 | 6 | 1 | 9 | 0.798 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 4 | 0.682 | 4 / 6 (66.67%) | 6 | 0 | 2 | 0 | 5 | 6311.33 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | hybrid-rag-fast | hybrid-rag | canonical | 2 | 4.083 | 6 / 6 (100.00%) | 0 | 3.833 | 6 / 6 (100.00%) | 4.333 | 6 / 6 (100.00%) | 0.500 | 6 | 0 | 10 | 0.675 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 6 | 0.626 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 1 | 3227.83 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | vanilla-rag-wide | vanilla-rag | canonical | 2 | 4.083 | 6 / 6 (100.00%) | 0 | 3.833 | 6 / 6 (100.00%) | 4.333 | 6 / 6 (100.00%) | 0.500 | 6 | 0 | 6 | 0.846 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 7 | 0.620 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 4 | 5685.00 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | contextual-rag-high-recall | contextual-rag | canonical | 4 | 3.917 | 6 / 6 (100.00%) | 0 | 3.833 | 6 / 6 (100.00%) | 4.000 | 6 / 6 (100.00%) | 0.167 | 6 | 0 | 1 | 0.905 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 9 | 0.200 | 5 / 6 (83.33%) | 6 | 0 | 1 | 0 | 12 | 26844.50 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | graph-rag-wide | graph-rag | canonical | 5 | 3.833 | 6 / 6 (100.00%) | 0 | 3.667 | 6 / 6 (100.00%) | 4.000 | 6 / 6 (100.00%) | 0.333 | 6 | 3 | 8 | 0.833 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | N/A | N/A | N/A | 6 | 6 | 0 | 0 | 8 | 17524.17 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | lazy-graph-rag-fast | lazy-graph-rag | experimental | 5 | 3.833 | 6 / 6 (100.00%) | 0 | 3.667 | 6 / 6 (100.00%) | 4.000 | 6 / 6 (100.00%) | 0.333 | 6 | 1 | 3 | 0.882 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 8 | 0.595 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 2 | 3392.67 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | hybrid-rag-high-recall | hybrid-rag | canonical | 7 | 3.750 | 6 / 6 (100.00%) | 0 | 3.500 | 6 / 6 (100.00%) | 4.000 | 6 / 6 (100.00%) | 0.500 | 6 | 0 | 2 | 0.893 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 5 | 0.681 | 5 / 6 (83.33%) | 6 | 0 | 1 | 0 | 11 | 26029.67 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | lazy-graph-rag-balanced | lazy-graph-rag | experimental | 7 | 3.750 | 6 / 6 (100.00%) | 0 | 3.500 | 6 / 6 (100.00%) | 4.000 | 6 / 6 (100.00%) | 0.500 | 6 | 0 | 4 | 0.868 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 2 | 0.720 | 5 / 6 (83.33%) | 6 | 0 | 1 | 0 | 3 | 4911.50 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | graph-rag-fast | graph-rag | canonical | 9 | 3.667 | 6 / 6 (100.00%) | 0 | 3.500 | 6 / 6 (100.00%) | 3.833 | 6 / 6 (100.00%) | 0.333 | 6 | 0 | 7 | 0.844 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | N/A | N/A | N/A | 6 | 6 | 0 | 0 | 7 | 8759.67 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | graph-rag-rerank | graph-rag | canonical | 9 | 3.667 | 6 / 6 (100.00%) | 0 | 3.667 | 6 / 6 (100.00%) | 3.667 | 6 / 6 (100.00%) | 0.333 | 6 | 0 | 5 | 0.847 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | N/A | N/A | N/A | 6 | 6 | 0 | 0 | 9 | 20866.33 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | n8n-adaptive-rag-default | n8n-adaptive-rag | canonical | 11 | 3.167 | 6 / 6 (100.00%) | 0 | 3.167 | 6 / 6 (100.00%) | 3.167 | 6 / 6 (100.00%) | 0.000 | 6 | 1 | 11 | 0.596 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 3 | 0.700 | 5 / 6 (83.33%) | 6 | 0 | 1 | 0 | 6 | 8537.00 | 6 | 6 | 0.00% | 0 | 0 |
| baseline_curated | 1 | agentic-rag-deeper | agentic-rag | canonical | 12 | 2.917 | 6 / 6 (100.00%) | 0 | 2.833 | 6 / 6 (100.00%) | 3.000 | 6 / 6 (100.00%) | 0.167 | 6 | 0 | 12 | 0.576 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 1 | 1.000 | 4 / 6 (66.67%) | 6 | 0 | 2 | 0 | 10 | 22713.50 | 6 | 6 | 0.00% | 0 | 0 |
| graph_native | 2 | hybrid-rag-high-recall | hybrid-rag | canonical | 1 | 4.188 | 8 / 8 (100.00%) | 0 | 3.875 | 8 / 8 (100.00%) | 4.500 | 8 / 8 (100.00%) | 0.625 | 8 | 1 | 1 | 0.855 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 7 | 0.327 | 5 / 8 (62.50%) | 8 | 0 | 3 | 0 | 7 | 10226.75 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | vanilla-rag-wide | vanilla-rag | canonical | 2 | 4.000 | 8 / 8 (100.00%) | 0 | 3.750 | 8 / 8 (100.00%) | 4.250 | 8 / 8 (100.00%) | 0.750 | 8 | 3 | 5 | 0.849 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 2 | 0.488 | 6 / 8 (75.00%) | 8 | 0 | 2 | 0 | 6 | 6094.25 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | lazy-graph-rag-balanced | lazy-graph-rag | experimental | 3 | 3.688 | 8 / 8 (100.00%) | 0 | 3.375 | 8 / 8 (100.00%) | 4.000 | 8 / 8 (100.00%) | 0.625 | 8 | 0 | 8 | 0.829 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 3 | 0.416 | 5 / 8 (62.50%) | 8 | 0 | 3 | 0 | 3 | 4731.38 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | lazy-graph-rag-wide | lazy-graph-rag | experimental | 3 | 3.688 | 8 / 8 (100.00%) | 0 | 3.625 | 8 / 8 (100.00%) | 3.750 | 8 / 8 (100.00%) | 0.375 | 8 | 1 | 6 | 0.843 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 4 | 0.401 | 5 / 8 (62.50%) | 8 | 0 | 3 | 0 | 4 | 5173.75 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | hybrid-rag-fast | hybrid-rag | canonical | 5 | 3.625 | 8 / 8 (100.00%) | 0 | 3.625 | 8 / 8 (100.00%) | 3.625 | 8 / 8 (100.00%) | 0.500 | 8 | 0 | 2 | 0.855 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 6 | 0.382 | 6 / 8 (75.00%) | 8 | 0 | 2 | 0 | 2 | 4380.50 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | lazy-graph-rag-fast | lazy-graph-rag | experimental | 5 | 3.625 | 8 / 8 (100.00%) | 0 | 3.375 | 8 / 8 (100.00%) | 3.875 | 8 / 8 (100.00%) | 0.500 | 8 | 1 | 7 | 0.833 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 8 | 0.314 | 6 / 8 (75.00%) | 8 | 0 | 2 | 0 | 1 | 3908.12 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | contextual-rag-high-recall | contextual-rag | canonical | 7 | 3.375 | 8 / 8 (100.00%) | 0 | 3.250 | 8 / 8 (100.00%) | 3.500 | 8 / 8 (100.00%) | 0.500 | 8 | 0 | 3 | 0.849 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 1 | 0.531 | 6 / 8 (75.00%) | 8 | 0 | 2 | 0 | 8 | 11707.50 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | graph-rag-wide | graph-rag | canonical | 8 | 2.625 | 8 / 8 (100.00%) | 0 | 2.500 | 8 / 8 (100.00%) | 2.750 | 8 / 8 (100.00%) | 0.500 | 8 | 1 | 10 | 0.816 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | N/A | N/A | N/A | 8 | 8 | 0 | 0 | 10 | 13365.50 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | graph-rag-rerank | graph-rag | canonical | 9 | 2.438 | 8 / 8 (100.00%) | 0 | 2.375 | 8 / 8 (100.00%) | 2.500 | 8 / 8 (100.00%) | 0.375 | 8 | 1 | 11 | 0.805 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | N/A | N/A | N/A | 8 | 8 | 0 | 0 | 12 | 14687.75 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | n8n-adaptive-rag-default | n8n-adaptive-rag | canonical | 10 | 2.375 | 8 / 8 (100.00%) | 0 | 2.250 | 8 / 8 (100.00%) | 2.500 | 8 / 8 (100.00%) | 0.250 | 8 | 0 | 12 | 0.740 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 9 | 0.000 | 1 / 8 (12.50%) | 8 | 0 | 7 | 0 | 5 | 5914.75 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | agentic-rag-deeper | agentic-rag | canonical | 11 | 2.250 | 8 / 8 (100.00%) | 0 | 2.250 | 8 / 8 (100.00%) | 2.250 | 8 / 8 (100.00%) | 0.250 | 8 | 0 | 4 | 0.849 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | 5 | 0.392 | 2 / 8 (25.00%) | 8 | 0 | 6 | 0 | 11 | 13527.00 | 8 | 8 | 0.00% | 0 | 0 |
| graph_native | 2 | graph-rag-fast | graph-rag | canonical | 12 | 2.062 | 8 / 8 (100.00%) | 0 | 1.875 | 8 / 8 (100.00%) | 2.250 | 8 / 8 (100.00%) | 0.625 | 8 | 0 | 9 | 0.821 | 8 / 8 (100.00%) | 8 | 0 | 0 | 0 | N/A | N/A | N/A | 8 | 8 | 0 | 0 | 9 | 11994.75 | 8 | 8 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | hybrid-rag-fast | hybrid-rag | canonical | 1 | 3.667 | 6 / 6 (100.00%) | 0 | 3.333 | 6 / 6 (100.00%) | 4.000 | 6 / 6 (100.00%) | 1.000 | 6 | 2 | 5 | 0.844 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 8 | 0.045 | 4 / 6 (66.67%) | 6 | 0 | 2 | 0 | 1 | 6984.50 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | contextual-rag-high-recall | contextual-rag | canonical | 2 | 3.500 | 6 / 6 (100.00%) | 0 | 3.167 | 6 / 6 (100.00%) | 3.833 | 6 / 6 (100.00%) | 0.667 | 6 | 0 | 10 | 0.811 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 2 | 0.357 | 1 / 6 (16.67%) | 6 | 0 | 5 | 0 | 12 | 40646.33 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | hybrid-rag-high-recall | hybrid-rag | canonical | 3 | 3.333 | 6 / 6 (100.00%) | 0 | 3.000 | 6 / 6 (100.00%) | 3.667 | 6 / 6 (100.00%) | 0.667 | 6 | 1 | 1 | 0.890 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 9 | 0.000 | 2 / 6 (33.33%) | 6 | 0 | 4 | 0 | 11 | 39889.17 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | lazy-graph-rag-fast | lazy-graph-rag | experimental | 3 | 3.333 | 6 / 6 (100.00%) | 0 | 2.833 | 6 / 6 (100.00%) | 3.833 | 6 / 6 (100.00%) | 1.333 | 6 | 1 | 7 | 0.827 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 3 | 0.240 | 2 / 6 (33.33%) | 6 | 0 | 4 | 0 | 2 | 7350.67 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | lazy-graph-rag-balanced | lazy-graph-rag | experimental | 5 | 3.000 | 6 / 6 (100.00%) | 0 | 2.667 | 6 / 6 (100.00%) | 3.333 | 6 / 6 (100.00%) | 0.667 | 6 | 1 | 3 | 0.873 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 6 | 0.121 | 3 / 6 (50.00%) | 6 | 0 | 3 | 0 | 3 | 7955.67 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | vanilla-rag-wide | vanilla-rag | canonical | 5 | 3.000 | 6 / 6 (100.00%) | 0 | 2.500 | 6 / 6 (100.00%) | 3.500 | 6 / 6 (100.00%) | 1.000 | 6 | 1 | 8 | 0.820 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 7 | 0.059 | 3 / 6 (50.00%) | 6 | 0 | 3 | 0 | 6 | 13370.50 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | lazy-graph-rag-wide | lazy-graph-rag | experimental | 7 | 2.917 | 6 / 6 (100.00%) | 0 | 2.833 | 6 / 6 (100.00%) | 3.000 | 6 / 6 (100.00%) | 0.833 | 6 | 0 | 11 | 0.805 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 5 | 0.208 | 2 / 6 (33.33%) | 6 | 0 | 4 | 0 | 5 | 10781.33 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | n8n-adaptive-rag-default | n8n-adaptive-rag | canonical | 8 | 2.583 | 6 / 6 (100.00%) | 0 | 2.000 | 6 / 6 (100.00%) | 3.167 | 6 / 6 (100.00%) | 1.167 | 6 | 0 | 2 | 0.877 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 4 | 0.224 | 3 / 6 (50.00%) | 6 | 0 | 3 | 0 | 4 | 10773.17 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | agentic-rag-deeper | agentic-rag | canonical | 9 | 2.417 | 6 / 6 (100.00%) | 0 | 2.000 | 6 / 6 (100.00%) | 2.833 | 6 / 6 (100.00%) | 0.833 | 6 | 0 | 12 | 0.709 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | 1 | 0.531 | 2 / 6 (33.33%) | 6 | 0 | 4 | 0 | 8 | 17243.00 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | graph-rag-rerank | graph-rag | canonical | 9 | 2.417 | 6 / 6 (100.00%) | 0 | 2.167 | 6 / 6 (100.00%) | 2.667 | 6 / 6 (100.00%) | 0.500 | 6 | 0 | 9 | 0.820 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | N/A | N/A | N/A | 6 | 6 | 0 | 0 | 10 | 31947.50 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | graph-rag-wide | graph-rag | canonical | 11 | 2.167 | 6 / 6 (100.00%) | 0 | 1.833 | 6 / 6 (100.00%) | 2.500 | 6 / 6 (100.00%) | 0.667 | 6 | 0 | 4 | 0.864 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | N/A | N/A | N/A | 6 | 6 | 0 | 0 | 9 | 26667.67 | 6 | 6 | 0.00% | 0 | 0 |
| cyber_threat_intel | 7 | graph-rag-fast | graph-rag | canonical | 12 | 1.917 | 6 / 6 (100.00%) | 0 | 1.667 | 6 / 6 (100.00%) | 2.167 | 6 / 6 (100.00%) | 0.500 | 6 | 0 | 6 | 0.841 | 6 / 6 (100.00%) | 6 | 0 | 0 | 0 | N/A | N/A | N/A | 6 | 6 | 0 | 0 | 7 | 16581.83 | 6 | 6 | 0.00% | 0 | 0 |