PRODUCTIZED · FIXED PRICE

RAG-in-a-Box.

A complete retrieval-augmented system on your documents, deployed in 5 weeks. Up to 10,000 docs indexed, citations on every answer, admin panel your team operates. Stop your team from saying "I know we documented this somewhere."

$15,000
5-week delivery
50% deposit · Balance on delivery
10,000 docs indexed 30-day post-launch support Citations on every answer
// PRODUCT SHEET
RAG-in-a-Box
SKU · DLS-RIB-V1
// Duration 5 weeks
// Documents Up to 10,000
// Generation Claude Sonnet 4.6
// Interface Web-based UI
// Support 30 days included
// IP Ownership Yours on payment
// Investment $15,000 USD

Four numbers that
define this scope.

10K

Documents indexed

// Maximum corpus size

5wks

Delivery timeline

// Discovery to production

30day

Support window

// Post-launch coverage

1

Refresh pipeline

// Auto-reindex schedule

Is this the right starting point?

RAG-in-a-Box is designed for one specific situation: a team with valuable documents that nobody can find. If that's you, this is the fastest, most predictable way to unlock them. If your situation is different, we'll route you to a better fit.

// THIS IS FOR YOU IF

Your team can't find what you already wrote

  • You have a body of documents (under 10K) — wikis, policies, runbooks, help articles, contracts — that nobody can find when they need to
  • You're spending real time on "where is X documented?" or "what's our policy on Y?"
  • You want answers grounded in your actual documents, with citations — not generic AI guesses
  • One use case to start: internal Q&A, customer-facing doc search, or a compliance lookup tool
  • You can identify the documents and the people who own them in week 1
  • Budget for an initial AI build is real but lean — under $20K for a first deployment
// LOOK ELSEWHERE IF

Your situation needs something else

  • You have far more than 10,000 documents or millions of records — talk to us about Studios Growth or Enterprise
  • You need an agent that takes actions, not just answers questions — start with Claude Agent Builder
  • You haven't yet decided whether RAG is the right tool — start with AI Audit & Roadmap
  • You need on-prem or air-gapped deployment with strict data sovereignty — let's scope custom
  • You need permission-aware retrieval at enterprise SSO scale — Studios Growth
  • You're in healthcare/finance and need BAA or DPA signed first — Compliance Setup before deployment

Six things you get.
Fixed scope.

Every RAG-in-a-Box engagement ships the same six components. Here's exactly what shows up in your environment at the end of week 5.

// INCLUDED 01

Complete retrieval pipeline

Document ingestion, intelligent chunking, embedding generation, vector database setup, hybrid search (semantic + keyword), and reranking. The full engineering stack for production-grade RAG — built and tuned to your content.

// INCLUDED 02

Up to 10,000 documents indexed

PDFs, Word, Markdown, HTML, PowerPoint, plain text, Google Workspace docs, Notion, Confluence. Scanned documents get OCR'd. Tables and structured content preserved with their context.

// INCLUDED 03

Web-based query interface

Branded query UI deployed at your domain. Conversational interface with citations on every answer, follow-up question handling, and conversation history. Works on desktop and mobile.

// INCLUDED 04

Admin panel

Your team's control surface — upload new documents, trigger reindexing, view what's been searched, see retrieval quality scores, and manage user access. No engineering required to operate.

// INCLUDED 05

Automated refresh pipeline

Nightly reindexing of changed documents from your source systems (Google Drive, Notion, Confluence, SharePoint). Stale content is the #1 RAG quality killer — solved at install.

// INCLUDED 06

30-day post-launch support

Direct Slack/email access to the engineer who built your system for 30 days after launch. Retrieval tuning, edge case handling, and "this query isn't working" fixes — all covered.

From document audit to queryable knowledge
in 25 working days.

Every RAG-in-a-Box engagement runs on the same five-week cadence. Weekly milestones, weekly check-ins, no scope drift.

01
// DISCOVERY

Document audit

Kickoff, document inventory, ownership mapping, eval question set agreed. We map your knowledge before we index anything.

// END-OF-WEEK
Doc audit + eval set
02
// ARCHITECTURE

Chunking strategy

Vector DB selection, embedding model choice, chunking & retrieval architecture, reranking layer designed.

// END-OF-WEEK
Architecture approved
03
// INDEX

Pipeline built

Ingestion runs, all documents indexed, retrieval tested against eval set. First retrieval quality scores delivered.

// END-OF-WEEK
Full corpus indexed
04
// INTERFACE

Query UI + admin

Branded query interface built. Admin panel deployed. Citation display, conversation flow, edge cases handled. Friday demo.

// END-OF-WEEK
Live UI in sandbox
05
// LAUNCH

Production handoff

Production deployment, refresh pipeline live, runbook delivered, training session recorded. 30-day support begins.

// END-OF-WEEK
In production + handoff

Six artifacts. All yours.

The concrete files and access you receive at the end of week 5. Yours to operate, modify, and own — no vendor lock-in.

{ }

Full source code

Ingestion pipeline, retrieval layer, query UI, admin panel — all in your repository. Yours on full payment.

Deployed system

Production-ready RAG system on your domain, in your cloud or ours, with your credentials and admin access.

Indexed knowledge base

Your documents fully indexed in the vector database, with the refresh pipeline running nightly.

📋

Eval scorecard (.xlsx)

The Q&A test set we built in week 1, with retrieval quality scores for each — your baseline for ongoing tuning.

📖

Operations runbook

How to add documents, monitor retrieval quality, tune the system, and handle edge cases. Troubleshooting guide included.

Training session video

Recorded 45-60 minute walkthrough showing your team how to operate the admin panel, watch anytime, share internally.

Six inputs to make this land in 5 weeks.

RAG-in-a-Box hits its timeline because we lock scope tightly and you bring the right inputs on day one. Here's what we need from you.

// 01
Your documents

Access to up to 10,000 documents — either as a folder, a source-system API (Notion, Google Drive, Confluence), or a bulk export.

// 02
A project sponsor

One stakeholder with authority to approve at week 1 (doc audit), week 2 (architecture), and week 4 (UI demo).

// 03
20-30 example queries with answers

The questions your team or users actually ask, with the correct answer for each. This becomes your eval set.

// 04
Source-system access

API credentials or admin access to whichever source systems host your documents — for the refresh pipeline.

// 05
API & infrastructure budget

Anthropic API key, embeddings provider account, and vector DB (we'll recommend in week 2). Pass-through costs only.

// 06
50% deposit

$7,500 deposit on SOW signature kicks off week 1. Balance due on production handoff at end of week 5.

Things buyers ask
before clicking.

The honest answers to the questions buyers ask in the intake call. If yours isn't here, ask us when we talk.

Three options. (1) Start with your most-queried 10K and add more later — most clients find 5,000-8,000 high-value documents serve 80%+ of real queries. (2) Add additional document tiers at $2,500 per 5,000 docs as a change order. (3) Move to a Studios Growth engagement ($35K-$75K) for unlimited document volume, permission-aware retrieval, and multi-source pipelines. The 10K limit isn't a technical ceiling — it's the scope boundary for fixed-price delivery.
All standard formats: PDF (including scanned with OCR), Word (.docx), PowerPoint (.pptx), Excel (.xlsx with description), Markdown, HTML, plain text, RTF. Source-system native formats: Google Docs/Sheets/Slides, Notion pages, Confluence pages, SharePoint documents, Box files, Dropbox files. Code repositories with documentation. For unusual formats (proprietary CAD files with descriptions, legal-specific formats), we'll confirm during the intake call. If your content is structured data (database records, API logs), that's a different pattern and may need Studios Growth scoping.
Well-built RAG dramatically reduces hallucinations but doesn't eliminate them. Three controls built in by default: (1) the model is instructed to cite sources and abstain when no relevant content is retrieved; (2) the eval suite we build in week 1 scores every retrieval against your test cases — you see the actual accuracy number on your content before launch; (3) low-confidence answers display a "low retrieval confidence" notice instead of guessing. Typical accuracy on well-curated corpora: 88-95% on the eval set. Accuracy depends heavily on document quality and clarity — we'll be honest in week 1 if your content needs cleanup before indexing will work well.
Your documents live in your chosen vector database (Pinecone, pgvector, Qdrant, Weaviate — recommended in week 2 based on your data and infrastructure). At query time, only the small retrieved chunks relevant to the specific question are sent to the model API for synthesis — never your full document set. We default to providers with enterprise terms guaranteeing no training on API inputs (Anthropic, OpenAI Enterprise). For maximum data sovereignty, the entire stack including generation can run on open-source models in your infrastructure — that may push timeline by 1-2 weeks but stays within the productized price. Details in our Privacy Policy.
Three components, all pass-through (we never mark up): (1) Vector DB hosting — typically $70-$400/month for under 10K docs depending on provider. (2) Embedding regeneration — usually under $100/month for steady-state operation, more during initial indexing. (3) Generation API costs — heavily query-volume dependent, typically $100-$800/month for internal teams under 100 users. We model expected costs in week 2 and build a cost dashboard so you can monitor in real time. Optional Studios retainer ($5K-$25K/month) covers ongoing optimization and feature additions for clients who want active management.
Yes, that's the refresh pipeline. We default to nightly reindexing of changed documents from your source systems. For systems with webhook support (Notion, Google Drive, Confluence), we can move to near-real-time updates — typically 5-15 minute lag. Stale content is the #1 RAG quality killer, so the refresh pipeline is core scope, not an upsell. You can also manually trigger reindex from the admin panel whenever you want.
Yes — customer-facing doc search is one of the most common RAG-in-a-Box deployments. Two notes for customer-facing use: (1) we recommend an additional eval pass with adversarial queries (people will ask the system everything, including out-of-scope questions) — we include this for customer-facing deployments at no extra cost. (2) We strongly recommend adding rate-limiting and basic monitoring for cost control on consumer traffic — included by default. For embed-in-your-site widgets (a chatbot bubble), we deliver the embed code as part of the week 5 handoff.
RAG-in-a-Box includes single-tier authentication — logged-in users see all indexed documents, public users see only what you mark public. For more granular permission-aware retrieval — where different users see different documents based on their existing access in source systems — that's part of Studios Growth scope (typically $35K-$75K). If you have a mix of public help-center content and internal-only content, two RAG-in-a-Box deployments (one public, one authenticated) is often cheaper than one Growth engagement.
You own the custom code, custom configurations, and indexed knowledge base — assigned to you on full payment. We retain ownership of our methodologies and reusable RAG patterns (our "Background IP"), but those are licensed to you perpetually as embodied in your deliverables. The source code is in your repo. The vector DB is on your account. There's no lock-in. If you want to take this to another vendor or your internal team in six months, everything you need is in the handoff package. Full IP terms in our Terms of Service §8.
The 50% deposit is non-refundable once we start week 1 work. If we materially fail to deliver against the eval set we agreed on in week 1, we work to fix it under our limited services warranty — re-perform or correct, that's covered in Terms §11. If you terminate without cause mid-engagement, you owe for work performed through termination. We're confident enough in fixed-scope productized delivery that we publish prices and timelines on the website — the warranty terms reflect that confidence.

From document chaos to
queryable knowledge in 35 days.

Book a 30-minute intake call. We'll confirm fit, scope the corpus, and send an SOW within 48 hours if it's a match.

Fixed price · 10K documents indexed · 30-day post-launch support included