Use case · Catalogue discovery

Find the datasets that actually answer the question.

A demonstration for the Department of Climate Change, Energy, the Environment and Water (DCCEEW). The evidence behind an answer is spread across documents, tables, catalogue records and reference data. First it has to be brought within reach. Then it has to be connected. Then someone, or something, has to decide whether it is enough to answer.

A landscape of monitoring stations under bands of atmosphere. A highlighted path links a pollutant to a place and then to the dataset that measures it.

01 · The gap

The model can't see your catalogue.

A language model knows a lot about air quality in general. It knows nothing about which datasets you hold, at what breakdown, from which publisher, or how current they are. One request shows how many kinds of evidence a good answer needs.

The requestAnnual PM2.5 by country, 2019 to 2024 — official sources only.

Evidence the request needs, its shape, and where it usually lives
What a good answer needsShapeWhere it usually lives
What PM2.5 means, and which pollutant group it belongs toproseMethodology documents
Which datasets hold PM2.5, for which countries and yearstablesCatalogue coverage tables
Who publishes each dataset, and when it was updatedrecordsCatalogue metadata
How countries roll up to continents, and years to rangesrelationshipsReference hierarchies
Whether this person may see the sourcepolicyAccess rules

02 · Reach

Two ways to bring evidence within reach.

Both end the same way: material added to the model's context window for this one request. Neither retrains the model. They differ in what they reach and who decides to reach for it.

RAG

Retrieval-augmented generation · a technique

Index the documents ahead of time. When a request arrives, search the index and add the best-matching passages or images to the prompt.

Here, it suits
The methodology documents: what PM2.5 is and how it is measured.
Who decides
Usually your code, before the model starts.
Fresh as
The last time the index was rebuilt.

MCP

Model Context Protocol · an open standard

A standard way for an AI application to connect to servers that offer tools, resources and prompts. The model asks for a tool; the application makes the call and returns the result [2].

Here, it suits
The catalogue itself: query coverage, read metadata, fetch the current record.
Who decides
For tools, the model, while it works. The application can still ask a person first.
Fresh as
The system behind the server.

They combine. A search_methodology tool can be an MCP tool with a RAG index behind it. RAG is a way of finding things; MCP is a way of connecting to things.

The shape of the data matters as much as where it lives.

RAGData shapeA well-built MCP tool
strong
What it was built for: chunk, embed, rank, cite the passage.
Textmethodologies, reportsgood
A search tool returns passages — usually from a retrieval index behind the server.
good
Multimodal embeddings or captions make charts, maps and scans findable.
Imagesmaps, charts, scansgood
Can return the original image with its record; finding the right one still takes search.
weak
Chunks cut rows from headers, embeddings blur numbers, and a handful of passages can't sum, filter or join.
Tablescoverage, measurementsstrong
The tool runs the query — exact filters, joins and totals — and returns only the rows needed.
mixed
Flattening records to text loses field names: fine for "similar datasets", poor for "official, updated since 2024".
Recordsmetadata, JSON, logsstrong
Reads fields by name, filters by value, and can return typed results checked against a schema [2].

Ratings are qualitative judgements from common practice, not benchmark scores.

  • A connection, not a data engine.MCP's advantage comes from the server behind it. A server that returns a text dump is no better than a chunk.
  • Only what's exposed.No aggregate tool, no total. Someone builds, secures and maintains each tool.
  • Wrong queries still run.If the model writes the filter, a plausible mistake returns a confident, well-formatted wrong answer.
  • Results fill the window.Large results must be filtered or aggregated on the server, or they crowd out everything else [3].

RAG and MCP bring evidence within reach. Neither checks that the pieces fit together, cover the whole request, or are enough to answer it.

03 · Before any request

Raw storage has to be sifted before either route works.

Connect Amazon S3, Azure Blob Storage or Google Cloud Storage and you get millions of objects with no schema: PDFs beside CSV exports, JSON logs, scanned forms and maps. A bucket server can list and fetch objects, but a file name doesn't say what is inside or whether it matters. Embed everything and the CSVs are indexed as prose, noise and all.

Production pattern · not part of this demo

Object storage Buckets and containers

S3, Blob Storage, Cloud Storage: PDFs, CSVs, JSON, logs, scans, maps

Sift · bounded decisions A Jev-style decision engine
  • Is it relevant to this domain?
  • What shape is it?
  • May it be shown, and to whom?
  • Where should it go?

Each call returns a confidence. Below the threshold, the answer is review.

Text and imagesIndex for RAG
Tables and recordsQueryable store behind MCP tools
Links between themGraph: pollutant, place, dataset, publisher, permission
Unsure or sensitiveReview queue, or skip
Sifting is millions of narrow yes-or-no and pick-one calls, not open-ended reading, so it suits a decision engine better than a bigger model. This happens when data arrives, not when a question is asked. The demo below starts after this point: its catalogue graph is already built.

04 · Connect

Access isn't understanding. The graph shows how the pieces fit.

Each route returns a fragment. The graph holds the relationships between fragments, so a coverage question can be checked instead of guessed.

RAG returns a passage"PM2.5: fine particulate matter, 2.5 micrometres or less…"

True, but it says nothing about which datasets hold it.

An MCP tool returns a recordEuropean Air Emissions Reporting · official · updated 2025

Accurate, but it doesn't say whether it covers this request.

The graph adds the linksPM2.5 → by country, by year → held by this dataset → from this publisher

That's what turns two fragments into one answer.

The worked request, as a path through the catalogue graph

  1. Request
  2. asks for
  3. PM2.5
  4. broken down by
  5. Country · Year
  6. covered by
  7. European Air Emissions Reporting
  8. published by
  9. Official publisher

Because coverage is stored as relationships, the graph can answer: does every requested country and year have a source? If no single dataset covers them, can two combine? Does a country-level source roll up to continents? The path above is what the demo returns for this request.

Graph traversal is not the same as Microsoft GraphRAG

This demo follows explicit, typed links in a catalogue graph. Microsoft GraphRAG is a specific method: it extracts an entity graph from text, groups it into communities and summarises them, and improved some whole-corpus questions over a vector-RAG baseline [4]. Both use graphs; they solve different problems.

05 · Decide

Connected evidence still needs a decision.

Jev is the bounded decision layer. At each step it answers one narrow question from a fixed set of answers, with a confidence, and the pipeline acts only on answers that clear the threshold.

It isn't a retriever, a protocol or a database. It stores no facts. It decides what the agent may do with them.

  1. Can this catalogue answer the request?answer · clarify · reject
  2. Which pollutants does it mean?yes or no for each graph candidate
  3. At what breakdown?one level per dimension
  4. Which valid option best fits the preference?a ranked fit, only among options the graph proved
  5. Is any of these too uncertain?below threshold → review
  • Answer
  • Clarify
  • Reject
  • Review
  • No data

06 · Try it

Ask the DCCEEW catalogue.

This self-contained publication runs representative catalogue decisions in your browser. Start with an example, or write your own request.

Start with an example

Name a pollutant and how you want it broken down, for example by country and year.

In plain words, what makes one dataset better than another for you? For example: official sources only.

Loading the catalogue…

What this proves

What runs today, and what production adds.

Running in this demo

  • A catalogue graph of eight illustrative sources, held in memory.
  • Graph coverage checks, including combined sources and roll-ups.
  • Five bounded outcomes, with refusal and clarification before any search.
  • A traceable result: stages run, datasets matched, and why.
  • A self-contained browser demonstration deployable to any static host.

Added in production

  • A RAG index over methodology documents, maps and charts.
  • MCP servers in front of the catalogue and source systems, with delegated, short-lived access.
  • Ingestion and sifting from object storage.
  • A persistent graph database instead of the in-memory graph.
  • A Jev model with calibrated thresholds, and a human review queue.
RAG and MCP: where each one failsFailure modes paired by kind, for technical advisers.
KindRAG fails when…MCP fails when…
FindingThe right passage is never retrieved.The model picks the wrong tool or passes bad arguments.
AttentionToo many chunks dilute the window; text in the middle of long prompts is used less well [5].Long tool lists and big results crowd the window [3].
FreshnessThe index lags the source documents.The server reads a cache or a lagging copy.
AccessChunks reach people who may not see the source.Servers rely on long-lived keys instead of delegated access.
Blast radiusA wrong answer, stated confidently.A tool with side effects runs on a bad decision.
BothUntrusted text in a document or a tool result can carry instructions. MCP adds one more route: tool descriptions are read by the model too.

Jev's place in this table: it cannot make retrieval find the right passage, but it can refuse to answer, ask for detail or send the case to a person when the evidence doesn't support an answer.

SourcesLinks last checked 5 October 2026.
  1. Lewis, P. et al. "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." NeurIPS, 2020. arxiv.org/abs/2005.11401
  2. Model Context Protocol. Specification, revision 2026-07-28: server primitives and tools.
  3. Anthropic. "Code execution with MCP." 4 November 2025. anthropic.com — token cost of tool definitions and results.
  4. Edge, D. et al. "From Local to Global: A GraphRAG Approach to Query-Focused Summarization." 2024. arxiv.org/abs/2404.16130
  5. Liu, N. F. et al. "Lost in the Middle: How Language Models Use Long Contexts." TACL 12, 2024. aclanthology.org/2024.tacl-1.9
  6. The worked request, the catalogue and its sources are illustrative. The storage-sifting flow is an architecture pattern, not a feature of this demo.