Building an Internal AI Policy That People Will Actually Follow

Most internal AI policies fail on the day they are published. Not because they are wrong, but because they are unusable: eleven pages of restrictions, no approved alternative, and a tone that assumes the reader is trying to get away with something.

People respond to that policy the way people always respond to unusable rules. They ignore it, use a personal account, and stop asking questions. You end up with less visibility than before you wrote anything.

A policy that works has three properties. It is short enough to remember. It says yes to something specific. And it makes the compliant path the easy path. Here is how to build one.

Start with what you are actually protecting

AI policy” is too broad to be useful. In practice you are managing four distinct risks, and they need different treatment:

Risk What it looks like Fix type
Data leakage Customer data pasted into a consumer account Classification + approved tools
Accuracy Fabricated facts reaching a customer Review requirements
Compliance Regulated content processed improperly Explicit prohibitions
Cost Uncontrolled duplicate subscriptions Procurement path

 

Write the policy in those four sections. A single undifferentiated list of rules teaches nobody which ones matter, so people apply the same energy to formatting a blog post as to handling a customer record.

Rule one: classify data, not tools

The most common structural error is a policy organised around tools — “ChatGPT is banned, this other tool is approved”. Tools change monthly. The classification underneath does not.

Three tiers is enough. Four is too many; nobody remembers four.

Green — public or synthetic. Marketing copy, public documentation, generic code, anything already published or entirely made up. Use any approved tool freely. No review required.

Amber — internal but not sensitive. Meeting notes, internal drafts, non-sensitive analysis, general strategy discussion. Approved tools only, on company accounts, never personal ones.

Red — regulated or confidential. Customer PII, health or financial records, credentials, unreleased financials, anything under contractual confidentiality. Never enters a general-purpose assistant. If there is an approved path — a specific reviewed platform with the right contractual terms — name it explicitly.

Then give three concrete examples per tier, drawn from your own business. Abstract definitions get argued about. Examples get followed.

Rule two: publish an approved stack, not a ban list

A ban list has an infinite tail. New tools launch weekly, and your list is out of date the day it ships. Worse, a ban list gives people nothing to do except find something not yet on it.

Publish instead:

  • The default stack. One or two named tools that cover the common cases, purchased centrally, with accounts everyone already has.
  • The specialist exceptions. Named tools for particular functions, with the owner listed.
  • The request path. How to get something new approved, with a realistic turnaround. “Message the owner, expect an answer in two working days” is a request path. “Submit to the committee” is a deterrent.

The default stack is doing the real work here. If the approved tool is genuinely good, shadow AI mostly stops on its own — people were not being difficult, they were routing around a gap.

That is also the strongest argument for a multi-model default rather than a single-vendor one. When the approved account covers several labs’ models, the most common reason to go outside policy — “the approved model refused this, or handled it badly” — disappears. Platforms like Perspective AI exist for exactly this shape of problem: one account, one set of retention terms, several model families behind it. The governance benefit is separate from the cost benefit and is often larger.

Rule three: match review requirements to consequence

Blanket review requirements are ignored because most AI output does not need reviewing. Tie the requirement to what happens if the output is wrong.

  • Customer-facing, published, or contractual → human review before release, always. No exceptions, no matter how good the draft looks.
  • Internal decision input → verify any factual claim you would not personally vouch for. Numbers, citations, dates, quotes.
  • Draft, brainstorm, or scratch work → no review. This is most AI usage and requiring review for it destroys the policy’s credibility for the cases that matter.

Name the failure mode in the policy itself, because people who understand why comply more reliably than people who are simply told. Language models generate plausible text, and plausible text includes plausible fabrications: citations that do not exist, statistics with no source, quotes nobody said. This is a property of the technology, not a bug to be patched. Say so.

Rule four: give guidance on model choice, briefly

Most policies stop at “use approved tools” and leave people to guess which model. That guarantees everyone uses whichever one is the default, for everything, including the tasks it handles worst.

Two or three lines is enough:

  • Long documents and careful reasoning → the model with the largest context window and strongest reasoning scores
  • Code → whichever model currently leads on coding benchmarks
  • Current events and citations → a model with live web search
  • Anything sensitive → the private or no-retention mode, if your platform has one

Point to a maintained reference rather than hardcoding model names, because names change. A public composite AI model leaderboard that averages coding, maths, reasoning, and human-preference benchmarks is a reasonable thing to link, with the caveat that any such ranking is a dated snapshot rather than a permanent verdict — check when it was last updated before treating a position as current.

Rule five: make the cost path obvious

Cost belongs in the AI policy, not only in the finance one, because uncontrolled subscriptions are the mechanism by which sensitive data ends up in unmonitored accounts. The two problems have the same root.

State plainly:

  • Which subscriptions are centrally provided, and how to get access
  • The expense threshold below which people can buy their own tools, if any
  • Who owns AI spend and approves exceptions
  • That personal accounts must not be used for Amber or Red data, under any circumstance

The last one is the important line. It is also the one most likely to be broken quietly if the approved tool is worse than the free one people already have.

The one-page format that works

Fit the whole thing on a single page:

  1. What this covers — two sentences
  2. Data tiers — the table, with three examples each
  3. Approved tools — the current list, with owners
  4. Review rules — three lines by consequence level
  5. Prohibited, always — a short, genuinely short list
  6. How to ask — a name and a channel
  7. Last reviewed — a date

Anything longer gets skimmed once and never reopened. If your legal team needs a longer version, keep it as an appendix and let the one-pager be what circulates.

What to review quarterly

AI policy goes stale faster than most governance documents. Four things to re-check every quarter:

  • Approved tool list — has anything been superseded, deprecated, or repriced?
  • Vendor retention terms — these change, and they change quietly
  • Incidents — what actually went wrong, and does a rule need to change?
  • Shadow usage — are people routing around the policy, and where?

That last one is a diagnostic, not an enforcement exercise. Shadow AI is a signal that the approved path has a gap. Find the gap and close it; that fixes the behaviour permanently, while enforcement fixes it until people get busy again.

Frequently asked questions

Should we ban AI tools outright? Almost never, and it rarely works. Bans push usage onto personal accounts and unmonitored devices, which is strictly worse than governed usage. The narrow exception is genuinely regulated content where a specific legal obligation applies.

How long should an internal AI policy be? One page for the operative version. Longer documents get referenced during onboarding and never again.

Who should own the AI policy? Someone who uses AI daily, with sign-off from legal or security. A policy written entirely by people who do not use the tools will be either unusable or unenforceable, sometimes both.

How do you detect shadow AI usage? Expense reports, browser telemetry if you already run it, and simply asking with an explicit no-blame framing. The last one surfaces more than the first two, and it costs nothing.

The point

A policy is not a legal shield, it is an operating instruction. It works when the compliant path is the easy path — a good default tool, a short set of memorable rules, and a real answer when someone needs something different.

Write the one page. Buy the tool that makes it followable. Review it in ninety days.

 

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *