Contents
- Before Build vs. Buy: Your Data Is the Actual Decision
- What data readiness actually means
- The readiness test (4 binary questions)
- Filter 1: Can a Managed Platform Index What You Need?
- What platforms index well
- Where platforms break
- Filter 2: Does Compliance Kill Cloud-Native Vendors?
- The air-gap question
- Compliance that doesn’t require air-gap
- Filter 3: Does the Math Favor Buy or Build?
- Real vendor prices
- The crossover math
- 10X Framework application
- The Scenarios Where We’d Tell You Not to Hire Us
- Vendor Reference Guide (Not a Review)
- Build vs Buy RAG: The Decision in One Place
- Frequently Asked Questions
- What’s the difference between RAG and an AI chatbot? Do I need both?
- Glean says they support custom document indexing. Why would I still need a custom build?
- How long does a custom RAG build actually take?
- What does data readiness actually cost to fix?
- Can I start with a managed platform and switch to custom later?
We tell roughly one in three RAG prospects to buy a platform instead of hiring us. That’s math, not modesty. Glean at 80 seats with a clean SaaS corpus beats a $60K custom build on every axis. Bedrock Knowledge Bases for mid-volume search over clean S3 buckets costs $500/month and takes a week. This post is for people who don’t yet know which category fits. If you want to understand the five architectural decisions that drive cost and complexity differences between build and buy, that post breaks down each decision point.
Five sequential filters follow. Each eliminates a category of company. What’s left is where hiring a studio like ours is the right call.
Before Build vs. Buy: Your Data Is the Actual Decision
Filter 0 sits upstream of every vendor question, and most RAG buyer guides skip it.
What data readiness actually means
Picking the vector database feels like the first decision. It’s closer to the last. On a messy corpus, the opening stretch of the project goes to taxonomy, deduplication, and access mapping: which system holds the authoritative copy of a document, who is allowed to see it, and what you do when two versions of the same manual disagree.
None of that is RAG work. You pay for it whether you buy Glean, wire up Azure AI Search, or hire someone to build custom. The platform choice doesn’t remove the bill. It changes who sends it.
Data prep routinely eats 40 to 50% of total RAG project budget on messy corpora. We’ve seen it consistently. Ask every vendor to price data readiness before they quote you a platform.
The readiness test (4 binary questions)
Answer these before talking to any vendor:
- Can you point to a single authoritative source for every document category you plan to index?
- Is access control enforced at the document level, not just the folder level?
- Does your corpus contain fewer than three distinct format types?
- Is the corpus updated on a predictable schedule by a defined owner?
Any “no” means your first spend is data remediation, not a platform. Typical remediation: 4 to 12 weeks, $15K to $40K.
Filter 1: Can a Managed Platform Index What You Need?
Data is clean. The question is whether off-the-shelf can take it from there.
What platforms index well
Glean, Vectara, Bedrock Knowledge Bases, and Azure AI Search handle SaaS content, structured web pages, standard PDFs in cloud storage, and Microsoft ecosystem documents. Glean in particular isn’t RAG over proprietary domain docs. It’s smarter search across your SaaS stack (Slack, Confluence, Jira, Google Workspace, Salesforce). If that’s your problem, Glean solves it.
Vectara sits one level deeper: an API-first managed RAG service for teams that want production RAG fast without custom retrieval logic. Bedrock and Azure AI Search are infrastructure-layer options for teams already inside AWS or Microsoft.
Where platforms break
The platforms break in predictable places: custom metadata schemas, domain-specific chunking (legal clause-level, engineering spec section-level), multi-LLM routing, anomaly detection, uncommon language combinations, and anything touching an air-gap. For teams evaluating agent platforms beyond RAG, OpenAI Agent Builder hits the same kind of disqualifier wall. Map your requirements before committing.
They also stack, and the stacking is what decides the question. One awkward requirement rarely justifies a build. Four at once usually does.
Take a corpus that spans two languages, needs clause-level chunking because the answer lives in a subsection rather than a page, routes different question types to different models because no single model handles them all well, and flags anomalies in what it retrieves. Each of those is a customization the managed platforms either don’t expose or expose too shallowly to be useful. Together there’s nothing left on the shelf to buy.
Run the buy evaluation anyway, and run it before anyone drafts a build proposal. Most companies would rather buy than build. So would we. A build only earns the money once the managed options have been priced and knocked out on requirements, in writing.
If your requirements map cleanly to Glean, Vectara, or Bedrock, go buy. Otherwise, Filter 2.
Filter 2: Does Compliance Kill Cloud-Native Vendors?
Filter 1 disqualifies on requirements. Filter 2 disqualifies on where the data is legally and physically allowed to sit.
The air-gap question
Every managed RAG platform assumes your documents can traverse the internet and sit on vendor infrastructure. Defense contractors, critical infrastructure operators, and companies under strict data residency regimes can’t make that assumption at any price.
Air-gap is the hard version of this problem. Data residency is a contract clause. An air-gap is a physical fact: no outbound connection at all, sometimes because the system runs somewhere reliable internet doesn’t reach. That pushes you to an open-weight model on hardware you control, a vector store inside the perimeter, and an evaluation loop that never calls out to anything. No managed vendor survives that requirement, and no budget brings one back.
Say the trade-off out loud before someone signs off on it. You give up frontier model quality, and you own the ops burden for the life of the system.
Compliance that doesn’t require air-gap
GDPR, HIPAA, and SOC 2 don’t automatically disqualify managed vendors. Most serious ones offer compliant options. What they add is contractual overhead: BAAs, DPAs, pen test reports, and procurement cycles measured in quarters. Get the compliance matrix before the pricing.
Filter 3: Does the Math Favor Buy or Build?
Assume Filters 0 through 2 passed. Now it’s a pricing question.
Real vendor prices
| Vendor | Pricing model | Annual cost | Best fit |
|---|---|---|---|
| Glean | Per seat + AI add-on | $97K+ median ACV; Y1 TCO $350K to $480K mid-size | Broad internal SaaS search, 200+ users |
| Vectara | Tier pricing | $100K (Small) / $250K (Medium) / $500K (Large) | Production RAG fast, clean docs, standard retrieval |
| Bedrock Knowledge Bases | Pay-per-use | $345/mo min OpenSearch Serverless or $50-80/mo Aurora pgvector; under $500/mo all-in for sub-50K queries | AWS-native, clean corpus, one engineer available |
| Azure AI Search | Per SKU + LLM | S1 around $250/mo; S3 around $2,750/mo; agentic retrieval billed via LLM separately | Microsoft shops, SharePoint-heavy |
| ChatGPT Enterprise | Per seat | $30/user/mo, 150-seat minimum, roughly $54K ACV | Knowledge worker productivity, NOT production RAG |
| Writer | Enterprise AI platform | Mid-market $75K to $250K; enterprise $500K+ | Large orgs wanting unified AI platform with brand controls |
Note: Mendable pivoted to Firecrawl and is now a web scraping API. If a vendor list still includes Mendable, it’s stale.
The crossover math
A studio build lands at $40K to $60K plus $1K to $3K/month running. Year one: $52K to $96K all-in. Year three: $76K to $168K, because once the asset is yours, costs flatten.
Glean starts at $97K median ACV in year one and climbs with seat growth. By year three, mid-size deployments typically sit at $290K+ cumulative. The crossover:
- Under 100 users: Glean wins year one on speed.
- 150+ users over 18 months: custom pays off.
- 300+ users from day one: custom is cheaper from month one.
Bedrock and Azure at low query volume sit at $500 to $3,000/year. That’s frequently the right answer for a hybrid: managed vector store plus custom retrieval logic on top. You keep the infrastructure you don’t want to own and control the thin layer that differentiates the application. When that custom layer scales to multi-step reasoning and tool use, you’re in agentic RAG territory, where query decomposition, reflection loops, and multi-agent orchestration become viable architectural choices.
10X Framework application
We read RAG opportunities through the 10X Framework. Three lenses: impact (which number it moves, and how much of the work it takes on, 1 to 5), trust to earn (what a wrong answer costs, who decides, and how sensitive the data is, written in words from Little through Some, Real, A lot, to The most, because a number there invites averaging and strong impact should never cancel out a trust problem), and readiness (whether the data and a named owner already exist, 1 to 5). The combination lands in one of four tiers: build now, proof first, build with guardrails, or not yet.
Apply it to a 200-user enterprise deploying custom RAG over proprietary technical manuals:
- Impact: 4/5. It carries a real slice of specialized knowledge work, and the hours are already reported somewhere.
- Trust to earn: Real. Custom builds carry genuine exposure from messy data, drift and multi-model complexity. Gartner projected 30% of generative AI PoCs abandoned by end of 2025.
- Readiness: 4/5. Clean data (Filter 0 passed) and a named owner both assumed.
Real trust to earn against strong impact and readiness lands in Tier 3, build with guardrails. That is a build verdict with conditions attached. Budget 20 to 30% of the engagement for evaluation pipelines, rollback infrastructure and structured human review.
The Scenarios Where We’d Tell You Not to Hire Us
Buy Glean if your search problem is “find things across our SaaS stack” rather than “query proprietary docs with custom logic,” you’re under 100 users, you need it live in 30 days, and IT can run the platform without custom retrieval needs.
Use Bedrock or Azure AI Search directly if you’re already on AWS or Azure, your corpus is clean and structured in cloud storage, an engineer can spend two to four weeks on the pipeline, and query volume is predictable enough for pay-per-use.
Fix data first if your corpus lives across three or more systems with no unified access, nobody owns document maintenance, or someone has promised that “just connecting SharePoint” will work when it won’t.
Hire us if your requirements include any of these: custom retrieval logic, air-gap or on-prem deployment, domain-specific chunking, multilingual corpora, anomaly detection on retrieval output, or a hybrid of managed infrastructure with a custom layer on top. One of those on its own is worth a conversation. Two or more and the managed platforms are already out of the running, whatever their sales engineers tell you.
Vendor Reference Guide (Not a Review)
Glean wins for 200+ user enterprises with broad SaaS search needs and no ML team. Base $45 to $50/user/month, AI add-on $15/user/month, seat minimum 100 to 250. Closed architecture; you can’t customize chunking, retrieval, or model.
Vectara wins for startups and mid-market teams that need production RAG fast on clean documents. Tiers run $100K / $250K / $500K per year. Custom reranking and unusual document types push you to higher tiers. Skip it if you need deep retrieval customization.
Bedrock Knowledge Bases wins for AWS-native teams with clean structured docs. Connect S3, Bedrock manages vector storage, embedding, and retrieval against Claude, Titan, or Llama. Trade-offs are limited chunking customization and AWS lock-in. Often the right answer under 50K queries/month at sub-$500/month; full pricing here.
Azure AI Search wins for Microsoft shops. Hybrid BM25 and vector search, deep SharePoint, Teams, and M365 integration. Brilliant for SharePoint content, awkward the moment a source system sits outside Microsoft.
ChatGPT Enterprise wins for knowledge worker productivity at $30/user/month, 150-seat minimum. File uploads give you basic RAG but not custom chunking or retrieval strategy. Corpus size is limited. Don’t confuse it with a production RAG platform.
Writer is the AI platform play, not pure RAG. Mid-market $75K/year, enterprise $500K+, bundling content generation and brand governance around a Knowledge Graph. Overkill if you only need RAG.
Build vs Buy RAG: The Decision in One Place
| Filter | Pass condition | Outcome |
|---|---|---|
| 0. Data Readiness | 4/4 binary checks pass | Proceed to Filter 1 |
| 1. Platform Fit | Requirements match Glean/Vectara/Bedrock indexing | Buy a managed platform |
| 2. Compliance | Air-gap not required, cloud-native acceptable | Continue on cloud track |
| 3. Math | Seat count, query volume, timeline favor custom | Build custom or hybrid |
| 4. Opt-out scenarios | Any “buy” condition above applies | Buy; skip the build |
If you’ve run these filters and still aren’t sure which row you land in, that’s what a scoping call is for. We’ll tell you which category fits, including if the answer is “buy Glean.” Book a call, or the chatbot build-vs-buy companion.
Frequently Asked Questions
What’s the difference between RAG and an AI chatbot? Do I need both?
RAG is a retrieval architecture. A chatbot is an interface. You can have either without the other. A RAG system can feed a batch pipeline with no conversational UI. A chatbot can run on a closed fine-tuned model with no retrieval. Most enterprise use cases pair them because the interface is conversational and the source of truth is your documents. For mechanics, read the architecture breakdown.
Glean says they support custom document indexing. Why would I still need a custom build?
Glean’s custom connector handles standard-format documents from cloud storage. What it doesn’t do: multi-LLM routing, custom reranking logic, anomaly detection layers on the retrieval output, or domain-specific chunking (clause-level legal, section-level engineering specs). If your use case needs any of those, the custom connector isn’t the answer.
How long does a custom RAG build actually take?
Six to ten weeks for a clean corpus and a single source. Around 14 weeks once indexing spans a hybrid of systems, SharePoint plus something that isn’t SharePoint. Around 20 weeks when retrieval itself gets complicated: multi-model routing, agentic query handling, an anomaly check on the output. Timeline is driven almost entirely by data complexity, not LLM choice.
What does data readiness actually cost to fix?
Four to 12 weeks of work. $15K to $40K in direct cost. On messy corpora, data prep eats 40 to 50% of total project budget before any retrieval code gets written. Plan for it explicitly or watch it consume your timeline silently.
Can I start with a managed platform and switch to custom later?
Yes, and migration cost is real. Bedrock and Azure AI Search give a cleaner migration path than Glean because they expose APIs your custom layer can consume or replace. Glean’s closed architecture makes a switch closer to a rebuild. If the odds of outgrowing the platform feel high, pick the vendor that lets you own the thin custom layer sooner.