Custom AI
AI that answers from your files, not from the internet
We build assistants, knowledge bases, and document tools grounded in your actual contracts, past proposals, and service history. Every answer cites the source document, so your team can check it in one click.
Built for
- Firms whose expertise lives in 10 years of documents nobody can search
- Proposal and capture teams rewriting the same content every cycle
- Support teams answering the same 40 questions forever
- Professional services firms where one senior person is the bottleneck
- Anyone who tried ChatGPT on company work and got a confident wrong answer
Custom AI builds run $15k–$75k. A focused internal assistant over a defined document set is usually $15k–$30k, while multi-source knowledge bases with permissions, or proposal tooling with human-in-the-loop review, run $40k–$75k. Ongoing tuning, evals, and support are $2,500–$12,000/mo. Claude API usage is billed to your account and typically runs $50–$600/mo for internal tools at normal team sizes.
The problem
Generic AI doesn’t know your business. That’s not a small gap.
Ask a general chatbot about your warranty terms and it will invent something reasonable-sounding. It has never seen your contract. The knowledge you actually need is in a 2019 proposal, a signed scope of work, and the head of estimating’s memory, none of which are on the internet.
So people go back to asking around. The new hire messages the senior estimator, who answers the same question for the eleventh time. The capture team searches SharePoint, gives up, and rewrites a past performance narrative that already existed in better form two folders away. The knowledge exists. It just isn’t findable at the moment someone needs it.
What it buys you
The outcomes we hold ourselves to.
- Every answer shows its work
- Responses cite the document, section, and date they came from, with a link, so your team verifies in one click instead of trusting a paragraph. If the system doesn’t have a source, it says it doesn’t know. That’s a setting we don’t let anyone turn off.
- The senior person stops being the search engine
- When ten years of scopes, warranties, and job history are searchable in plain English, the new hire gets the answer in 30 seconds without interrupting anyone. Your best people go back to the work only they can do.
- Proposals start at a draft, not a blank page
- Assembled from your own approved past performance, resumes, and technical language, matched to the section you’re answering. Teams that spent 3–4 days on a first draft typically get to a same-day draft they then edit. The editing is the value; the retyping never was.
- Documents read themselves
- Contracts, POs, subcontractor agreements, and inspection reports get parsed for the fields you care about, checked against your rules, and flagged when something’s off. A human reviews exceptions instead of reading all 200 pages.
- Support answers get consistent
- The AI drafts from your real policies and past resolved tickets, and a person sends it. Same answer whoever’s on shift, and the deflection is on the questions that were never worth a human anyway.
- It stays current without a project
- New documents get indexed automatically from SharePoint, Google Drive, or wherever they live. The knowledge base doesn’t rot six months after launch, which is how most of them die.
Deliverables
What you actually receive.
Written, handed over, and yours to keep, whether or not we work together again.
Internal assistant
- Chat interface in the tools your team already opens: Slack, Teams, or a simple web app
- Grounded on your documents with citations on every answer
- Permission-aware: people see answers only from documents they could already open
- “I don’t know” as a real, tested response path
- Question logs so you can see what people actually ask
Knowledge base and retrieval
- Document ingestion from SharePoint, Google Drive, Dropbox, or a file share
- Chunking and embedding tuned to your document types, not a generic default
- Vector search on Postgres with pgvector, plus keyword search for part numbers and contract IDs
- Metadata filters: date, project, client, contract vehicle
- Automatic reindexing as documents change
Document and proposal tools
- Extraction pipelines for contracts, POs, invoices, and inspection reports
- Rules-based validation with exceptions routed to a named human
- Proposal drafting from your approved past performance and resume library
- Compliance matrix generation from an RFP’s Section L and M
- Meeting summarization with owners and due dates pulled out
The engineering underneath
- Built on Claude via API, running in your cloud account
- Evaluation set of 50–200 real questions with known-good answers, run before every change
- Cost controls and per-user rate limits so nobody gets a surprise bill
- Full logging of prompts, sources, and responses for audit
- Source code and infrastructure handed over in your repositories
Process
How the engagement runs.
No surprises, no scope drift you didn’t agree to. Each phase ends with something you can read or use.
- 01
Pick the question that’s worth answering
We find the specific, repeated question that costs you real time: “what did we quote this client in 2022,” “does this scope cover the warranty.” Narrow beats broad every time. A system that answers one question well gets used; one that answers everything vaguely gets abandoned.
Week 1 - 02
Look at the documents honestly
We sample your real files and tell you what’s retrievable. Scanned PDFs from 2011 with no OCR are a different problem than clean Word docs. You get a straight read on coverage before anyone builds anything.
Weeks 1–2 - 03
Build the evaluation set first
Before the system exists, we write down 50–200 real questions and the answers your experts would give. That’s the scoreboard. Without it, “it seems good” is the only review you’ll ever get, and it’s worthless.
Week 2 - 04
Build, measure, tighten
Retrieval, then prompting, then the interface. We run the eval set after every meaningful change and show you the score, and where it fails you see the failures, including the ones we haven’t fixed yet.
Weeks 3–8 - 05
Put it in front of five people
A small group uses it on real work for two to three weeks while we watch the logs and fix what breaks. Then it opens up. Nothing goes wide until it’s survived contact with people who don’t care about the demo.
Weeks 8–12
Examples
The kind of work this turns into.
- Past performance library for a mid-size gov contractor
- Twelve years of proposals and CPARS across two SharePoint sites, indexed with contract vehicle, agency, and NAICS as filters. Capture managers ask “what have we done for DLA in logistics IT since 2020” and get cited excerpts. The first draft starts from real approved language instead of a blank template.
- Scope and warranty assistant for a commercial contractor
- Every signed scope, change order, and warranty from the last decade, searchable in Slack. When a customer calls arguing about coverage, the PM has the exact clause and the document it’s in before the call ends. The estimator stopped being on-call for history questions.
- Compliance matrix from an RFP
- Upload the solicitation, get Section L and M parsed into a matrix with every requirement, its cross-reference, and a first-pass owner assignment. It turns a two-day exercise with a highlighter into an afternoon of review, and it doesn’t miss a line because someone was on hour six.
- Support drafting for a services firm
- Incoming tickets get a drafted reply built from real policy docs and past resolved tickets, with sources attached, sitting in the agent’s queue. The agent edits and sends. Answers stopped depending on who picked up the ticket.
- Meeting notes that produce actions
- Recordings transcribed and summarized into decisions, owners, and dates, then pushed into the CRM against the right deal. The follow-up exists before everyone forgets who said they’d do it.
Questions
Before you ask.
How do I know this isn’t just a chatbot toy?
Because we score it and show you the score. We write 50–200 real questions with known-good answers before we build, and the system is measured against them at every step, so you see the failures rather than a demo. If it can’t clear the bar on your own questions, we tell you and you don’t launch it, which is a better outcome than a toy nobody trusts.
What stops it from making things up?
It only answers from retrieved documents, and it cites them. If retrieval comes back empty, the response is “I don’t know” with a suggestion of who to ask, and that path is tested in the eval set like any other. The constraint is enforced in the architecture, and you can click through to the source on every answer to check it yourself.
Our documents are a mess: scans, duplicates, five versions of everything.
Normal. We sample them in week one and give you a straight read on what’s retrievable. Scans get OCR’d, duplicates get handled at indexing, and version conflicts get resolved by a rule you choose (usually most recent signed). If a chunk of your corpus genuinely isn’t usable, we tell you before you spend the money, not after.
Where does our data go? Some of it we can’t send anywhere.
It runs in your cloud account with your API keys. Anthropic doesn’t train on API data. For regulated work we can keep everything inside your boundary and document exactly what crosses which line. If a document set can’t leave your enclave, we architect around it or tell you it’s out of scope.
What happens when the AI models change?
The eval set is the answer. When a new model ships, we run your 200 questions against it and compare scores, so you get a number instead of a vendor blog post. Model choice is a config line. This is exactly why we build the scoreboard first.
Can our own developer maintain this after you’re gone?
Yes, and that’s the design. It’s standard Python or TypeScript, Postgres, and the Claude API, in your repositories, with the eval suite included. No proprietary framework. A competent developer can read it, and we do a handover walkthrough as part of the build.
Other services
AI Consulting
A written plan that says which three things to automate first, what each one costs, and what you get back.
Business Automation
Your systems talk to each other. Quotes, invoices, and follow-ups happen without anyone retyping anything.
Government & Enterprise
AI and automation that survive an ATO conversation, built for CUI boundaries, audit trails, and the review board.
Let’s find the work worth automating.
A 30-minute discovery call. We map where your time actually goes and tell you what’s worth automating, and what isn’t. No deck, no pressure. You leave with a written summary either way.