2.2 Quick Start¶
The whole showcase runs on Atlas (vendored as a
Git submodule at infra/) and comes up with one command.
1. Prerequisites¶
Atlas's requirements apply:
- Docker + Docker Compose 2.20.3 or newer (Atlas recommends 2.26.0+),
installed and running. The Compose top-level file merges fragments via the
native
include:directive. - The vendored
infra/submodule initialized: - Host tools
uvandpython3(Atlas's bootstrapper and the host-side corpus fetch use them). - An Atlas-supported LLM backend. The manifest commits
LLM_PROVIDER_SOURCE: auto, so each host resolves the best source (an existing host Ollama if installed, else a container) automatically — no per-run flag. - Disk / RAM / headroom for the
gen-ai-ragstack plus your chosen local models. The default local run uses Atlas'sqwen3.8:latestfor chat and all three LightRAG LLM roles, withnomic-embed-textfor embeddings — see Hardware Sizing for minimum and recommended profiles.
2. Bring It Up¶
This single script:
- Starts Atlas with durable
BASE_PORT: autoin the manifest, so Atlas resolves a completely free 110-port block below the OS dynamic/private range once and keeps it stable across restarts (persisted toinfra/.env). The project name israg-showcase(override withRAG_SHOWCASE_PROJECT_NAME). SetRAG_SHOWCASE_BASE_PORTto pin a specific block; Atlas rejects it if any port is occupied. - Selects
atlas.consumer.ymland runs Atlas's native headless env backfill, manifest-aware Compose validation, and consumer doctor. The manifest declares project/brand metadata, the env file, external Compose overlay, backend plugin root, nineteen LiteLLM aliases, Ollama sidecar, adaptive workflow, and RAG ingestion profiles without tracked Atlas modifications or a_usersymlink. - Starts the Atlas
gen-ai-ragstack with--no-tui --detach; Atlas applies therag-showcaseproject and brand metadata, waits on Compose health, and returns. On a fresh checkout, the initial bootstrap banner can retain Atlas artwork because Atlas renders it before applying the consumer manifest. The stack includes LightRAG, TEI reranker, Weaviate, Neo4j, n8n, Open WebUI, and LiteLLM. The showcase wrapper explicitly disables the hardware-dependent Docling source, so Atlas falls back to plain-text parsing and the selected profile's Chonkie recursive chunker. Atlas starts only the enabled service set and owns dependency and initial one-shot classification. - Proceeds after Atlas's detached health summary, then assembles the corpus
on the host (
corpus/fetch_corpus.py). If Atlas reports the known exited-zero one-shot race, the wrapper proceeds only when that exact log signature is present and a strict, provider-aware Docker-state check confirms every long-lived service is ready and every expected init service exited zero. - Waits for model readiness (embed + chat), submits the
showcase_defaultAtlas RAG ingestion job, waits on its machine-readable phase record, and then builds the contextual collection from Atlas-written plain chunks. Atlas compiles the consumer-declared LiteLLM aliases intoconfig.yamlbefore the proxy boots, so they are discoverable in/v1/modelsat startup with no consumer-side reconcile or restart. - Prints the Open WebUI URL.
First run downloads models
With local models, the first run may pull several GB, so it takes a while.
start-all.sh gates on model readiness — let it finish.
Running the graph approaches locally
LightRAG graph extraction is the heaviest local step and has host-specific
footguns (Ollama version skew, model-churn eviction, and an upstream
extract-runaway bug). If a run stalls at ingestion or graph-rag returns
nothing, see Running Graph Approaches Locally.
Then open the printed URL, start a multi-model chat, and select:
vanilla-rag, hybrid-rag, contextual-rag, graph-rag, agentic-rag,
n8n-adaptive-rag, lazy-graph-rag. One prompt fans out to all of them.
Stop everything with:
To customize Atlas-owned values without editing the submodule, copy the committed manifest and env file, point the manifest copy at the env copy, and select it:
cp atlas.consumer.yml atlas.consumer.local.yml
cp config/atlas.env.user .env.rag-showcase
# In atlas.consumer.local.yml set env.file: ./.env.rag-showcase, then edit both.
ATLAS_CONSUMER_MANIFEST="$PWD/atlas.consumer.local.yml" ./scripts/start-all.sh
Atlas's detached startup revalidates after applying the wrapper's source flags.
All compute sources are committed in atlas.consumer.yml (profile: dev +
env.values): LLM_PROVIDER_SOURCE: auto (host-adaptive), LightRAG in a container,
TEI on its CPU container, and Docling disabled. So a plain start is correct on every
host and needs no per-run source flag:
To pin or change a source, edit the manifest, or pass --llm-provider-source (etc.)
directly to infra/start.sh for a one-off — an operator flag wins that run.
ComfyUI is not required by the gen-ai-rag track, so omit its override unless an
enabled Atlas service actually needs it.
3. Corpus Note¶
For the full corpus (MultiHop-RAG + keyword docs), install the datasets library on
the host before running:
Without it, ingestion uses only the bundled keyword docs, so the thematic / multi-hop
demo queries have little to work with. See corpus/README.md.
The dataset ladder selects the matching manifest profile (baseline_curated,
graph_native, and so on) before the stack starts. Atlas owns discover, parse,
chunk, embed, base-vector write, LightRAG upload, drain, and phase status. The
showcase retains only contextual-blurb generation because that transform belongs to
the contextual-rag approach rather than generic infrastructure.
4. The n8n Workflow¶
The n8n-adaptive-rag workflow is checked in and declared in
atlas.consumer.yml. Atlas validates, namespaces, imports, activates, and probes it
— including activation with no N8N_API_KEY (Atlas #720), so the wrapper does no
manual publish or n8n restart; it only verifies the real production webhook before
reporting readiness. See n8n/README.md for
ownership, lifecycle, workflow shape, and tuning knobs.
5. Development and Testing¶
uv run pytest # unit suite (mocked I/O) + integration tests (skip without the stack)
uv run pytest backend_plugins # unit tests only
The unit tests mock all external I/O and run without the stack. The
tests/test_demo_matrix.py integration tests exercise the live stack and self-skip
when LiteLLM is unreachable. With a started stack they derive the published gateway
and master key from infra/.env automatically, so a plain uv run pytest tests
works; export LITELLM_BASE_URL / LITELLM_MASTER_KEY only to target a
non-default gateway:
To confirm the evaluation's Atlas-infra dependencies are up and in order — the LiteLLM aliases, Weaviate plus its ingested collections, LightRAG (health plus knowledge-graph population), the TEI reranker, n8n, and the required Ollama models (pulled, not just declared), and flags an Ollama client/server version skew — without running any ingestion, approach, or the LLM judge, run the read-only preflight against a started stack:
It exits non-zero if anything is missing or unreachable, so it is a cheap gate before an expensive matrix run.
Build these docs locally with:
Full environment-variable reference and troubleshooting live in the project README.