5.3 Live Comparison¶
A side-by-side comparison of the RAG approaches in this repo, run against a live
gen-ai-rag Atlas stack. The recorded 2026-07-17 run used a local workstation
with host Ollama; that is run metadata, not a repo requirement. See
hardware.md for hardware guidance without assuming one host shape.
- Run date: 2026-07-17
- Approaches compared: all 6 canonical approaches plus explicitly selected
experimental
lazy-graph-rag(7 base families), followed by all 12 named non-base flavor aliases as a separate tier. - Baseline corpus: 11-document curated corpus — a MultiHop-RAG subset plus
widget-error-codes.md. - Graph-native corpus: 10 committed relation-dense dossiers in
corpus/graph_native/. - Cyber corpus: 60 committed MITRE ATT&CK dossiers in
corpus/cyber_threat_intel/. - Queries: baseline prompts in
demo/queries.yaml, graph-native prompts indemo/graph_native_queries.yaml, and cyber prompts indemo/cyber_threat_intel_queries.yaml. - Harness:
compare/run_matrix.py,compare/summarize.py, andcompare/judge.py. Renewed runs write canonical evidence JSONL, deterministic evaluation JSON, and compatibility matrix/judgment JSON. Raw working outputs live under gitignoredcompare/results/; the ladder validated and committed matrix, judgments, canonical evidence JSONL, and deterministic evaluation JSON for each dataset underresults/. - Methodology: the full protocol, model-role map, judge-panel design, and
dataset-ladder process are documented in
evaluation-methodology.md.
These are historical Qwen3.6/Mistral run results. The active runtime now uses
qwen3.8:latest; model names below and in the raw artifacts remain unchanged
because they are provenance for the answers and scores actually recorded.
The experimental lazy graph family remains excluded from default, but its base
alias was selected explicitly for this run. Its implementation, cold/warm cache
measurements, and keep-experimental decision are documented in
lazy-graph-rag.md.
1. Headline¶
The renewed ladder completed its requested base-family and flavor cells across all three measured datasets without response errors or timeouts. Rankings shifted with the corpus, and the flavor tier changed the leading configuration on each measured rung. The generated base leaderboard and flavor leaderboard own the exact ranks, scores, coverage, failures, and latency values.
The current run depends on these integration fixes and operating choices:
- Atlas model-scoped
think:falsefor the configured Qwen reasoning model; - LightRAG role-specific EXTRACT/KEYWORD/QUERY models configured separately;
- LightRAG EXTRACT tuned to
max_async=1andtimeout=900; nomic-embed-textembeddings for graph ingestion;- LightRAG upload retry on HTTP 409 backpressure, with exact already-processed conflicts treated as idempotent during a resumed ingest;
- TEI rerank batching for both chunk and LightRAG candidates, capped to the reranker's 32-item client batch limit;
- Atlas-managed LightRAG query profiles for canonical, fast, wide, and rerank variants, all sharing one ingested graph.
2. Reproduce¶
./scripts/start-all.sh
export JUDGE_MODELS=judge-a,judge-b
uv run python compare/run_matrix.py
uv run python compare/judge.py
MATRIX_MODELS can restrict the approaches for a partial run:
MATRIX_FLAVORS expands named profiles from compare/flavors.yaml, which is the
benchmark-side companion to the backend flavors.yaml used for Open WebUI aliases:
The graph-native comparison uses the same harness with alternate input/output files:
export JUDGE_MODELS=judge-a,judge-b
MATRIX_QUERIES_FILE=demo/graph_native_queries.yaml \
MATRIX_DATASET_ID=graph_native \
MATRIX_RUN_ID=manual-graph-native \
MATRIX_RESULTS_FILE=graph_native_matrix.json \
MATRIX_CANONICAL_FILE=graph_native_evidence.jsonl \
MATRIX_SUMMARY_FILE=graph_native_evaluation.json \
uv run python compare/run_matrix.py
JUDGE_MATRIX_FILE=graph_native_matrix.json \
JUDGE_RESULTS_FILE=graph_native_judgments.json \
uv run python compare/judge.py
uv run python compare/summarize.py \
--rows compare/results/graph_native_evidence.jsonl \
--output compare/results/graph_native_evaluation.json \
--csv-output compare/results/graph_native_evaluation.csv \
--judgments compare/results/graph_native_judgments.json
For the dataset-by-dataset view, use the dataset manifest and report generator:
That report is committed at docs/dataset-complexity-report.md
and is organized by input dataset complexity rather than by vector/graph collection.
The committed 2026-07-17 ladder used the end-to-end runner so every measured dataset got a fresh ingest, LightRAG drain, matrix run, judge run, result snapshot, manifest update, and report regeneration:
JUDGE_MODELS=qwen3.6:latest,gemma4:31b \
uv run python scripts/run-dataset-ladder.py \
--date-stamp 2026-07-17 \
--dataset baseline_curated \
--dataset graph_native \
--dataset cyber_threat_intel \
--include-flavor-tier
3. The approaches¶
See the README for the entry table and
docs/approaches.md for exact internal steps, dependencies,
tuning variables, and measured behavior. In one line each:
vanilla-rag is dense top-k; hybrid-rag adds BM25 and TEI rerank;
contextual-rag retrieves context-prefixed chunks; graph-rag delegates to
LightRAG; agentic-rag runs a ReAct loop over vector and graph tools; and
n8n-adaptive-rag routes through the n8n workflow; experimental
lazy-graph-rag performs deterministic concept-graph expansion.
4. Environment¶
| Concern | This run |
|---|---|
| Hardware | Mac Studio M2 Ultra, 192 GB unified memory |
| Atlas | 2026-07-17 run pin c744467e (v0.1.0-438-gc744467e; current pin 7f2fcf2d), project rag-showcase; baseline/graph-native rows record pre-rerank-fix 2229fee9, cyber rows record c744467e |
| Ports | baseline and graph-native 64500-64609; cyber 22000-22109; both blocks were verified free at assignment |
| Provider | host Ollama selected through Atlas ollama-localhost; ComfyUI disabled as not applicable |
| Generation | local Ollama qwen3.6:latest, with Atlas-scoped think:false |
| LightRAG roles | EXTRACT mistral-small3.2:24b; KEYWORD/QUERY qwen3.6:latest |
| Embeddings | nomic-embed-text, host Ollama for LightRAG; LiteLLM embedding route for plugin vectors |
| Judges | qwen3.6:latest + gemma4:31b, local Ollama, think:false |
This run uses the current rag-showcase alignment to Atlas's public LIGHTRAG_*
role inputs. Current setup configures LightRAG through Atlas and lets Atlas/LiteLLM
decide whether model calls go to container Ollama, host Ollama, GPU container
Ollama, or another configured provider.
5. Findings¶
think:falseis mandatory for the Qwen reasoning model. With thinking enabled, extraction and generation calls spend time on hidden reasoning. Withthink:false, the same local model avoids unnecessary reasoning overhead for this workload. The setting is scoped per model, so it does not leak to unrelated models. The current baseline delegates the same model-scoped default to Atlas's model catalog.- Atlas now has a first-class host-Ollama source. The original run needed an
ad hoc LiteLLM alias to reach host Ollama. The updated Atlas submodule exposes
LLM_PROVIDER_SOURCE=ollama-localhost, so the repo no longer needs to assume any particular host hardware path. - Atlas now exposes LightRAG role-specific model settings. The current
submodule maps
LIGHTRAG_EXTRACT_*,LIGHTRAG_KEYWORD_*, andLIGHTRAG_QUERY_*inputs into LightRAG's native roles. Rag-showcase now sets those Atlas inputs instead of patching LightRAG runtime env directly. - LightRAG extraction works, but graph builds remain expensive. Fresh builds drained for all three corpora, including 60 cyber documents producing 66 chunks. The cyber extraction phase dominated ladder runtime even with the dedicated non-reasoning extraction model.
- LightRAG rerank is now compatible and measured. Atlas translates the LightRAG request to TEI and batches candidate lists over the 32-item service limit. The run exposed and fixed the missing batching case in Atlas #713 / #714, then discarded the pre-fix row and reran the complete flavor tier.
agentic-ragis still step-limited.MAX_STEPS=4is too low for several synthesis prompts; it does well on single-hop tool use and often stops early on multi-step tasks.
6. Reading the Current Results¶
The three measured datasets do not support a single best architecture. The baseline favors direct dense retrieval, the relation-dense dossiers reward the experimental lazy-graph path, and the cyber corpus favors context-prefixed retrieval. The flavor tier reinforces that query-time tuning is dataset-specific, rather than evidence that one base family is universally superior.
The complete leaderboards contain the generated base and flavor tables for every approach and metric, including coverage, failures, latency, Ragas, and judge-panel fields. The generated dataset complexity report retains the ladder and per-query views, while the artifact ledger identifies the committed evidence behind those tables.
7. Judgment Panel¶
The scoring pass used compare/judge.py, which evaluates stored
matrix answers after all approaches have already run. Its manifest selected local
Ollama models qwen3.6:latest and gemma4:31b, both called with temperature=0
and think:false.
The panel was chosen to keep evaluation local and repeatable while avoiding a
single-model judge. For each query, the harness anonymizes and deterministically
shuffles the approach answers, gives the judges the query-specific scoring
rationale from the query YAML, asks for 1-5 scores plus a best-answer letter, and
then aggregates mean score by approach with best-answer votes as the tiebreaker.
The judgment files in docs/results/ keep the per-judge scores, reasons, and
resolved judge backend. This published run used direct host Ollama. The checked-in
harness now defaults to the Atlas LiteLLM gateway and reads judge models, endpoint,
temperature, and optional thinking from compare/evaluation.yaml; environment
overrides can use another OpenAI-compatible provider without changing the runner.
8. Graph Findings¶
The renewed run shows that the graph path is technically healthy: LightRAG indexed the baseline, graph-native, and cyber corpora, drained extraction, and answered through the same LiteLLM/Open WebUI route as the other approaches.
The quality story is more nuanced. Default graph-rag remained operational
across the ladder, but its aggregate position varied by dataset. Atlas-managed
profiles materially changed both quality and latency, but no one graph profile
dominated every dataset. See the generated
base leaderboard
for the exact per-dataset comparisons.
The rerank-enabled profile is technically healthy but remains opt-in. Against
the other graph profiles, reranking changed quality and answer-relevancy tradeoffs
without a consistent benefit across datasets, while adding latency. The detailed
tradeoff table is in
approach-flavor-tuning.md.
The cyber corpus is the clearest warning against assuming that a graph-shaped
input automatically favors the LLM-extracted graph endpoint. The ATT&CK docs are
highly relational, but the judges favored contextual retrieval overall. Lazy
graph led the graph-native rung and makes no index-time LLM calls. It remains
experimental because co-occurrence
edges are untyped and its concept extractor is deliberately lightweight. See
lazy-graph-rag.md for the measured keep decision.
9. Caveats¶
- Bounded corpora: the scored run uses bounded corpora: 11 baseline docs, 10 graph-native dossiers, and 60 ATT&CK cyber dossiers. Larger graph builds are still a separate capacity test.
- Three rungs measured so far: the dataset complexity report still includes heavier candidates such as STaRK, OpenAlex, and GDELT, but their rankings remain pending until live matrix and judge snapshots are produced.
- Graph-native corpus is synthetic-curated: the documents are real-world dossiers with source links and explicit relationship bullets, designed to make graph structure available. This is a better graph test than the baseline subset, but it is still not a large natural enterprise corpus.
- Current local judges: scores are directional, not authoritative. Answers are shuffled and anonymized, but the 2026-07-17 panel used two local models. The renewed harness keeps that panel separate from Ragas and operational metrics.
- Faithfulness coverage varies: the local evaluator occasionally returned a
null faithfulness score. Those rows remain
partialwithscore_missing_or_null, are excluded from the mean, and appear in each ranking's evaluated/total coverage instead of becoming zero. - Cache effects: n8n and graph-rag include cache hits in some cells.
- Agentic cap:
MAX_STEPS=4materially limitsagentic-rag. - Profile sensitivity: fast, wide, and rerank profiles trade judge quality, answer relevancy, and latency differently by dataset; none is a universal replacement for the canonical profile.
10. Reversibility¶
qwen3.6-moewas a historical LiteLLM runtime alias used during the early live run before the Atlas submodule gained first-class host-Ollama support.- The recorded run used a local
models.yamlcompatibility layer forthink:false; the current baseline removed that layer because Atlas'sqwen3.6:latestcatalog entry owns the same scoped request default. - LightRAG role/query settings are Atlas inputs supplied by the parent-owned
config/atlas.env.userimported byatlas.consumer.yml. Use an alternateATLAS_CONSUMER_MANIFESTand env-file pair to experiment without editing the Atlas submodule.