Case Study

AtlasOS — a 60-agent litigation platform

Turning a raw case file into the work-product a litigation team actually uses

A litigation team starts with a pile of documents. Complaints, answers, contracts, discovery, deposition transcripts, expert reports, court orders, often thousands of pages, a lot of it scanned. Getting from that pile to an actual case strategy is weeks of associate time. AtlasOS is a system I built for a global top tier law firm that does the first pass. It runs about 60 specialised agents that read the case file and produce the documents a litigation team actually works from.

What it produces #

Not summaries. Real work product:

  • Proof maps. Claim by claim, element by element: what each side has to prove, what facts each side has, what evidence proves it, plus witness mapping and gaps in the evidence.
  • Case theory. Three competing one sentence theories of the case, ranked, with one recommended.
  • Trial themes. The persuasive throughline, with evidence anchors, the attack you should expect, and the rebuttal to it.
  • Deposition digests. A deposition is sworn witness testimony taken before trial. Every clip here gets coded to an issue with an exact page and line citation, plus a coverage report and a hot list.
  • Deposition outlines. An examination plan for each witness: vulnerabilities, admissions to go after, impeachment targets.
  • Motion drafts. Rule 12(b)(6) motions (arguing a complaint doesn't state a valid legal claim) and Rule 56 motions (summary judgment, asking a judge to decide the case without a trial), from the analysis through to full prose.
  • Opponent playbook. How the other side is likely to litigate this, laid out as a red team and blue team matrix.
  • Settlement strategy, jury instructions (the rules a judge reads to the jury before it deliberates), expert deposition outlines, and a capstone strategy memo.

The document pipeline #

Everything below happens once, at upload, so the actual generation later is fast.

OCR

A vision model converts the PDF to text. Large files get split into small page batches first, so the model never times out on a 600 page document. There are automatic retries built in, and a page by page fallback if a batch fails.

Structured extraction

Three specialist agents run over the text as it comes out of OCR, pulling people, events and dates, and legal citations into structured records that live separately from the raw text. That's why the cast of characters and the case timeline can show up instantly later. The work already happened at upload.

Chunk and embed

Text gets cut into roughly 1,500 token passages. It's sentence aware, so it never slices a sentence in half, and there's about 100 tokens of overlap so a fact sitting on a boundary doesn't get lost. Each passage gets embedded and stored in a vector index.

Classification

A separate pipeline files each document into the right folder in a fixed case taxonomy.

One thing that trips people up: two different things both get called "chunking" here. PDF chunking is an OCR problem, splitting a big file so the vision model can process it in parallel. Text chunking is a retrieval problem, cutting text into passages sized for the embedding model. The two have nothing to do with each other, and mixing them up is a fast way to misread the pipeline.

Every stored passage also carries metadata: which folder it came from, its page range, the people and dates mentioned in it, whether it has citations. That means retrieval can be filtered as well as semantic, things like "only pleadings" or "only passages that mention this witness."

The wave model, how 60 agents stay ordered #

The agents run in seven waves. Every agent inside a wave runs in parallel, and a wave doesn't start until the one before it has fully finished. That ordering is the whole point. It's what guarantees each agent's inputs actually exist before it runs.

W0

The case spine

The master structured record: parties, claims, defenses, jurisdiction, posture, deadlines. Almost everything downstream reads it.

W1

The foundational record

Cast of characters and timeline, the opponent playbook, response strategy.

W2

The order of proof

The element-by-element map of what each side must prove.

W2.5

Case theory

Crystallised before themes, so themes get built on an approved theory rather than a shifting one.

W3

Trial themes

The persuasive throughline, plus a tagging palette used by later agents.

W4

Narrative & trial prep

Narratives, jury charge, expert planning, demonstratives, early case assessment.

W5

Settlement & operations

Settlement architecture, litigation hold, team onboarding.

W6

The capstone memo

A strategy memo that integrates the core artifacts from every wave before it.

Inside a single agent #

The wave model above is about ordering across all 60 agents. Inside each individual agent, the shape is consistent too. Every agent that produces an analysis runs through the same eight steps. What changes from one agent to the next is what it depends on, what it reads, what it searches for, and what it's actually trying to produce. The skeleton underneath stays the same.

  1. Initialize. Pull in whatever the case already has: structured facts from earlier waves, plus an index of every document in the case (filename, folder, type), so the agent knows what exists before it has to read any of it.
  2. Plan. The agent looks at what it's supposed to produce and writes its own reading plan: which documents to open in full, and what to search for across everything else.
  3. Read and compress. It reads the documents from its own plan and compresses them into a focused summary before carrying them forward, so a 100 page complaint doesn't blow up the context window later on. Summarize first, synthesize after.
  4. Search. It runs both the fixed checklist queries and its own reasoned queries against the case's documents, to fill in anything the reading step didn't cover.
  5. Check coverage. Before writing anything, it checks whether what it's gathered actually covers every part of the output it owes. If something's missing, it says so instead of guessing.
  6. Fill gaps. If coverage came up short, it goes back for more: another document, another search, then checks coverage again. That loop is capped, so it can't run forever.
  7. Compress. Everything gathered gets condensed one more time into a tight evidence summary.
  8. Draft. The actual writing step. It produces the final output, checks that output against a quality bar for that specific work product, and files anything unresolved into the case's running list of open questions.

The dependencies, the prompt, and the output shape are different for every one of the 60 agents. The eight steps underneath them are not.

Hard vs. soft dependencies #

One of the earliest decisions I made ended up shaping everything downstream: almost every dependency between agents is soft.

An agent reads its upstream inputs if they exist. If they don't, it still runs, and it reports the gap instead of failing outright. Only a handful of dependencies are hard. Case theory genuinely will not run without the case spine, the foundational record, and the order of proof. The summary judgment tool won't run without a proof map and a deposition record.

That tradeoff was deliberate. If every dependency were hard, the system would be brittle and constantly stuck waiting on something. Soft dependencies mean it almost never blocks a user, but the failure mode shifts from an error to something quieter: run something too early and you get a thinner result instead of a crash. Making those gaps visible in the output was how I dealt with that.

The chain is resilient in the same way. If one agent errors out, it gets logged and the run keeps going instead of taking the other fifty nine down with it.

Agentic retrieval #

This is the part I find most interesting, honestly. Retrieval queries come from three places, and they sit on a spectrum from hardcoded to reasoned.

Fixed queries are the checklist: things a competent lawyer always checks for a given work product. The litigation hold agent always looks for spoliation (destroying or altering evidence) and preservation rules. The response strategy agent always looks for service of process (formal notice that someone is being sued) and jurisdiction (whether this court actually has authority over the case). These are the floor, not the whole picture.

Agentic queries are where it gets interesting. Most agents have a planning step where they see an inventory of every document in the case (filename, folder, type) and from that they write their own reading plan and their own search queries for whatever they're about to produce. The agent decides what to open and read in full, and what to search for across everything else.

That means two cases running the exact same agent can pull completely different passages. The retrieval is reasoned, not scripted. A trial themes run ends up generating queries about that specific case's hot documents. An early case assessment generates queries about that case's own pivot facts.

Web research is a third channel, separate from the other two, used by only three agents, and only for public information like a judge's background or how a venue (the court where a case is heard) has ruled in similar verdicts before.

Not guessing #

In a legal setting, a confident wrong answer is worse than admitting you don't know something. A few design choices come straight out of that:

  • If the classifier can't confidently decide what a document is, it gets left unfiled and flagged instead of guessed at.
  • If an agent's inputs are missing, it reports the gap instead of filling it in.
  • Deposition digests carry exact page and line citations, and rough or uncertified transcripts get detected, with their citations flagged as approximate.
  • The motion drafter leaves citation placeholders instead of making up case cites. This is probably the single most important guardrail in the whole system.
  • Agents write into running "open questions" and "master schedule" lists, so uncertainty turns into a visible work item instead of a silent hole.

Per-case isolation #

Every case gets its own private partition in the vector index. A search in one case can never surface another case's text, and deleting a case, or even a single document, cleanly removes its vectors too. For a firm that sometimes has clients on opposite sides of different cases, that's not optional.

Stack #

Python FastAPI LangGraph Pinecone Vision model (OCR) Hosted embeddings