Back to TokenGrill
Deployment

The Local Model Engine

What actually runs on the appliance: a RAG-enabled, legal-customised engine built on open-weight Llama 3 and Mistral models, grounded on your own documents.

What the engine is

  • Open-weight Llama 3 and Mistral models, tuned on legal-domain tasks: extraction, clause classification, review and drafting.
  • A retrieval layer that indexes your own matter documents locally and grounds every answer in them.
  • A citation-grounding pass that flags claims your indexed documents do not support.
  • Local statutory and court-rule databases, refreshed by signed monthly updates.

How work is executed

A request is broken into steps on the appliance — extract, analyse, draft — and each step runs against your indexed documents on the local engine. Steps that can run in parallel do. Every step is attributed in the audit log, and no step has a path off the box.

Keeping it current

  • Your subscription includes regular pushes of faster, smarter open-weight model builds.
  • Statutory and court-rule databases sync monthly so citations reflect current law.
  • Zero-day security patches arrive as outbound-only signed updates; nothing is sent back.

Need this in writing for your review?

We share the underlying documentation with security and risk teams under NDA, and will walk your counsel through anything on this page.

Request documentation