Government

Compliance-aware AI for government contractors: a plain-English starting point

What you can and can’t put into a public model when you hold CUI. NIST 800-171, CMMC, and FedRAMP explained without the acronym fog, plus the questions to ask any vendor.

Nate Daniels9 min read

If you hold a DoD contract with the 7012 clause in it, your staff are already pasting things into chatbots. Not because they’re careless, but because it’s useful and nobody told them where the line is.

This is the briefing we give government contractors before we build anything. It isn’t legal advice and it’s no substitute for your compliance officer. It’s meant to get you to the right questions faster than reading the source documents cold.

The three acronyms and what they actually govern

These get used interchangeably, and they aren’t the same thing.

NIST SP 800-171 is the control set. It lists what you must do to protect Controlled Unclassified Information (CUI) in your own systems: 110 controls across 14 families. If DFARS 252.204-7012 is in your contract, this is your obligation, and it has been for years.

CMMC is the verification mechanism. It’s how DoD checks that you actually did the 800-171 things you said you did in your self-assessment. Level 1 covers basic FAR safeguarding for federal contract information. Level 2 maps to 800-171 and is where most contractors handling CUI land. Level 3 adds controls from 800-172 and applies to a small set of programs. The important shift is that self-attestation is no longer sufficient at Level 2 for many contracts. A third party assesses you.

FedRAMP is about cloud services. If a cloud service processes, stores, or transmits your CUI on your behalf, 800-171 control 3.1.20 and the 7012 clause push you toward requiring that service to meet FedRAMP Moderate baseline or equivalent. This is the control that decides almost every AI question you have.

What you actually cannot do

Let’s be concrete, because the abstract version leads to bad guesses.

  • You can’t paste CUI into the consumer version of a chatbot. Not ChatGPT free, not the personal-tier subscription your engineer expensed, not the browser extension summarizing your documents. None of these are authorized to hold CUI, and in some cases the terms permit training on your inputs.
  • You can’t put a technical data package, an ITAR-controlled drawing, or a CUI-marked requirements document into a general-purpose consumer service. This is the one that causes actual incidents.
  • You can’t assume an enterprise tier fixes it. Enterprise plans typically add a no-training commitment and better admin controls, which is necessary but nowhere near sufficient. What matters is whether the service carries a FedRAMP Moderate authorization or a documented equivalent, and whether your contract accepts that. Most consumer-facing enterprise tiers don’t.
  • You can’t rely on a vendor’s marketing page saying “secure” or “compliant.” Ask for the authorization, its scope, and the FedRAMP Marketplace listing. If they can’t produce it, the answer is no.

What you can do

The list is longer than people expect, and this is where most of the actual value sits.

First: everything that touches no CUI at all. Marketing copy, general research, internal scheduling, drafting a non-sensitive email, summarizing a public solicitation from SAM.gov. Public RFPs are public. There is nothing stopping you using a commercial model to help you understand a public document. A surprising amount of proposal work lives here.

Second: authorized government cloud environments. Azure OpenAI in Azure Government, AWS Bedrock in GovCloud, and similar offerings exist precisely for this. They carry the authorizations, they run inside boundaries you can document, and they’re the standard answer for CUI workloads. They cost more and they lag the commercial models by a bit. That’s the trade.

Third: models running inside your own accredited boundary. If you already have an enclave assessed for 800-171, running an open-weight model on hardware inside it means no data leaves and nothing crosses a boundary, so the boundary question never comes up. The trade is that you own the operational burden and the model is weaker than frontier commercial options.

Fourth — and this is the underrated one — deterministic automation. Most of what government contractors want from AI is document routing, form population, deadline tracking, and compliance checklisting. None of that needs a model. Plain automation inside your existing boundary solves it with no new compliance surface at all. We suggest this more often than we suggest AI, and clients are usually relieved.

The question that decides everything: where is the boundary?

Before any tooling conversation, you need a documented answer to this: which systems are in scope for CUI, and where exactly does the line sit?

Many contractors have never drawn this properly. They have a vague sense that “the engineering share drive is the sensitive one” and no diagram. Then someone connects an AI note-taker to the company calendar, the note-taker joins a program review, and CUI is now sitting in a transcript on a vendor’s cloud. Nobody decided that. It just happened.

Do this first, before AI is on the agenda:

  1. Inventory where CUI lives today. Actually look. It’s in email, in the file share, in the ticketing system, and probably in a folder on someone’s laptop.
  2. Draw the boundary: what’s in scope, what’s out, and what crosses. One diagram, one page.
  3. Inventory the SaaS tools already touching in-scope systems. Include browser extensions and anything with calendar or inbox access. This is where the surprises are.
  4. Classify AI use cases against the boundary before evaluating vendors. Inside-boundary and outside-boundary are different projects with different budgets.
  5. Write an acceptable use policy that names specific tools. “Don’t put sensitive data into AI” is a wish, not a policy. “Approved: Azure OpenAI in our GovCloud tenant. Not approved: any consumer chatbot, any browser extension, any meeting transcription tool not on this list” is a policy.

Questions to ask any AI vendor

Send these in writing. Vagueness in the answer is itself an answer.

  • Do you hold a FedRAMP authorization? At what baseline, what is the scope, and can I see the Marketplace listing?
  • Where is my data physically processed and stored, by region? Name the regions.
  • Is my data used to train or improve any model, under any tier, under any circumstance? Get this in the contract, not the FAQ.
  • What’s your data retention period, and can I set it to zero?
  • Are subcontractors or subprocessors involved in processing, and are they in scope of your authorization?
  • Will you sign a DFARS 7012 flowdown? Will you report cyber incidents to DoD within 72 hours as required?
  • Do you support US persons–only access to support systems that could reach my data?

A vendor that answers these crisply is a vendor that has done this before. A vendor that responds with a security overview PDF and enthusiasm has not.

A realistic sequence

For a mid-sized contractor starting from zero, this is roughly how we’d phase it.

Start with the boundary diagram and the SaaS inventory. Then publish the acceptable use policy, because shadow AI use is already happening and every week without a policy is a week of unmanaged risk. Then deploy AI on the outside-boundary work first: proposal research on public solicitations, marketing, internal non-CUI documentation. It builds the habit and delivers value while the harder conversation proceeds.

Only then take on the inside-boundary use case, in an authorized environment, with your compliance officer involved from the first meeting rather than at the end. And keep the deterministic automation option on the table throughout, because a fair share of what looked like AI work turns out to be a workflow that never needed a model.

The contractors doing this well aren’t the ones moving fastest. They’re the ones who drew the line first and then moved quickly inside it.

Let’s find the work worth automating.

A 30-minute discovery call. We map where your time actually goes and tell you what’s worth automating, and what isn’t. No deck, no pressure. You leave with a written summary either way.