Hackathon submission · Track C

Bridging the Gap

Advancing AI policy by comparing AI vendor disclosures to local government contracts

What can local governments meaningfully control, and where can local action fill gaps in broader AI policy?

Problem
Voluntarily disclosed AI vendor information is of little use if the claims are not reflected in actual contracts.
Solution
An analysis tool that helps procurement, legal, and administrative staff understand where contracts for AI-related services fall short.
Theory of change
The entities that build and deploy generative AI products, from frontier labs to genAI service providers, have little incentive to increase transparency and often insulate themselves from any and all liability. We seek to build on the efforts of the GovAI Coalition and their mission to harness the collective buying power of local government. While the Coalition has initiatives that encourage voluntary disclosures from AI service providers (Vendor AI FactSheet) and a repository of AI service contracts (AI Contract Hub), the two are not meaningfully connected.
Outputs
  • A master prompt public servants can use to compare a vendor’s AI FactSheet and proposed contract, highlighting where the two diverge
  • A sample case study report generated using real-world documents
  • An online version of the tool: mangrove-governance-hackathon.vercel.app

Background

Collectively, local governments wield more contract purchasing power than the federal government. Harnessing this potential could meaningfully impact AI policy. One initiative has taken the first step. The GovAI Coalition is a membership organization made up of over 3,000 public servants representing over 900 public agencies and governments.

Two resources developed by the Coalition are particularly relevant.

The first is an AI FactSheet, which the group refers to as an “AI Nutrition Label.” It is a standardized questionnaire that public agencies require commercial vendors to complete before procuring or deploying AI systems. It asks about items like model architecture, training data, accuracy and performance metrics, human oversight, cybersecurity compliance, and many more. Vendors are also encouraged to submit their FactSheets directly to an AI Registry.

The second resource is the AI Contract Hub. This is a contract repository created by the Coalition in partnership with Pavilion, a contract search platform. Many local governments lack specific technical, legal, and privacy expertise. The Hub was created to pool procurement resources across jurisdictions, increasing access, reuse, and learnings from already vetted and negotiated AI contracts.

Method

The Vendor AI FactSheet is based on the NIST AI Risk Management Framework, and we built directly on that foundation. We identified thirteen relevant FactSheet questions to check for in the contract analysis. The remaining nine fields, for things like Vendor Name or Overview, are documented but not scored.

For each of the criteria, the tool compares verbatim statements from both documents, then determines the severity of the mismatch by looking at a table. If the contract says nothing about a topic, that omission is tracked as a finding. Ensuring verbatim statements also helps the user later search and validate the findings.

Categorizing how claims can be validated

Reviewing the contract is only part of the evaluation a buyer should do. Some claims can be checked directly, and others cannot. We sorted every criterion, in advance, into one of three buckets:

Test

Run the product and watch what it does.

Access

Ask the vendor to let you look inside.

Warranty

Nothing done from the outside would ever reveal a false answer, so a contractual guarantee is the only protection.

The report closes with a section organized by these buckets, telling the reader what to test, what access to request, and what to get in writing.

How to use it

A couple of ways:

Case study

To test out the tool, we searched for an ideal use case and landed on Polimorphic’s resident-facing AI chatbot. Polimorphic had filed a FactSheet in the Coalition’s registry and executed a master services agreement.

Analysis takeaway

Polimorphic’s FactSheet is reasonably comprehensive, but little of it carried over to the contract. The common pattern in the analysis was “mentioned but not binding.”

10 high 1 medium 2 no finding

Featured findings

Training on data

The tool highlighted two clauses that might have otherwise escaped comparison. One promised the customer’s data would not be used to train models. The other gave the vendor broad discretion to use that same data, anonymized, to “improve its services” (which it did not define). This is the sort of point to be raised and clarified during a contract negotiation.

Lower bars

The fact sheet states 90% response accuracy, 1.5-second response times, and 99.9% uptime. But the contract only requires a “high level of system accuracy,” with no further detail. This matches an overall finding of specific claims turning vague in the contract.

Exceptions

The fact sheet says responses include references to source documents when appropriate. The contract takes this even further: it commits to citing sources and clearly labeling all AI-generated content. It is also a case of the prompt capturing that addenda to the contract (in this case a sales proposal) are also binding.

Limitations

There are a number of constraints we identified as we developed this project.

  1. Reliance on genAI judgment. The risks governments worry about in genAI deployment apply to our own process: inaccuracies and hallucinations. We grounded the analysis in the source material and included verbatim excerpts to enable verification.
  2. Reliance on a specific document pairing. While contracts are relatively easy to come by, not all vendors have a document like the AI FactSheet that outlines their claims. Additionally, it could be that there is a self-selection bias of sorts where the vendors most willing to voluntarily share documents like the FactSheet are on average better actors.
  3. Small sample size. We developed and tested against a single document pair, so the current run cannot validate the method against unseen documents. Additional testing would likely surface edge cases or other scenarios not anticipated in the current prompt version.

Appendix: the master prompt

Copy everything below into an AI assistant, attach the completed AI FactSheet and the contract, and send. The reply is the report.

Show the full prompt
You are helping a public agency check whether its contract for an artificial intelligence product delivers what the vendor disclosed about that product.

Two documents should be attached to this conversation. Read them, then work through the thirteen criteria listed near the end of this prompt, in order, and produce one report in the shape shown below.

## The documents

**FACT SHEET.** The vendor's completed AI FactSheet, filed to a shared registry run by the GovAI Coalition, a group of local, county, and state governments. It describes the product. It is a disclosure, not a binding agreement.

**CONTRACT.** The executed agreement. Treat everything in this file as contract text, including any vendor proposal, technical response, or marketing material bound into it, because such responses are commonly incorporated into the statement of work. Record where each statement appears. Do not decide which part governs if they disagree; that is a judgment for a lawyer.

Before you begin:

- Identify which attachment is which. The fact sheet is typically the shorter document, organized as questions and answers about one product; the contract is the executed agreement, usually longer, with numbered sections and signatures. Do not narrate this in the report; mention the documents themselves only if something is wrong.
- If fewer than two documents are attached, if either has no readable text (for example a scanned image without a text layer), or if you cannot tell which is which, stop and say so instead of producing a review.
- If you can see only excerpts of a document rather than the whole text — some tools retrieve fragments of large files instead of reading them end to end — say so prominently at the top of the report. A finding that something is absent is only trustworthy if you read the entire document.

## How the framework works

Each criterion names a topic, the fact sheet questions that usually cover it, an assurance type, and a contract expectation.

**Assurance type** — the only way a promise of that kind could ever be checked:

- **test** — A false claim shows up from outside by observing system behavior. The agency needs an acceptance standard and a test method.
- **access** — A false claim is catchable only with a staging environment, configuration view, or inspection right. The agency needs that access granted in the contract.
- **warranty** — No external observation reveals a false claim, at any effort. The agency needs a binding promise with liability attached.

**Contract expectation**:

- **required** — This criterion must appear as a contract term. Absence is a finding.
- **conditional** — This topic does not always call for a contract term. It is still evaluated, but counts for less when unaddressed. Which criteria are conditional is decided here, once, not per review.

For each criterion you will decide three things.

**Contract commitment**, about the contract only — whether the contract commits the vendor to anything on this topic:

- **absent** — The contract says nothing about this topic.
- **mentioned** — The contract raises the topic but does not commit the vendor to anything specific. No standard, no way to measure it, no consequence for failing, or it commits to something different from what this criterion needs, as described in its 'What good looks like' line.
- **binding** — The contract commits the vendor to something specific enough that a failure could be identified and acted on.

**Consistency**, about everything both documents say on the topic, taken together:

- **aligned** — The statements point the same way. One may be more detailed than another, but none contradicts the others.
- **tension** — The statements are hard to square, but could both be true, or one is vague enough to permit the other.
- **conflict** — The statements cannot both be true, or one takes back what another gives.

**Specificity**, scored 0 to 3 for each document separately:

- **0** — Unfalsifiable. No mechanism, no threshold, no standard named. ('industry-standard security', 'high level of accuracy')
- **1** — Names a mechanism but no measurable threshold. ('encryption in transit and at rest')
- **2** — Names a mechanism and a threshold, but no test method or remedy. ('99.9% uptime')
- **3** — Names a mechanism, a threshold, and how it is verified or enforced. ('AES-256 at rest, US-region tenant, SOC 2 Type II report provided annually')

## Severity

Severity is not your judgment. After you have decided contract commitment and consistency for a criterion, find the one row of this table matching that criterion's contract commitment, consistency, and contract expectation, and report the severity in that row. Every combination appears exactly once. Do not adjust the result for how important the topic feels.

| Contract commitment | Consistency | Contract expectation | Severity |
|---|---|---|---|
| absent | aligned | required | **high** |
| absent | aligned | conditional | **medium** |
| absent | tension | required | **high** |
| absent | tension | conditional | **medium** |
| absent | conflict | required | **critical** |
| absent | conflict | conditional | **high** |
| mentioned | aligned | required | **high** |
| mentioned | aligned | conditional | **low** |
| mentioned | tension | required | **high** |
| mentioned | tension | conditional | **medium** |
| mentioned | conflict | required | **critical** |
| mentioned | conflict | conditional | **high** |
| binding | aligned | required | **none** |
| binding | aligned | conditional | **none** |
| binding | tension | required | **medium** |
| binding | tension | conditional | **medium** |
| binding | conflict | required | **critical** |
| binding | conflict | conditional | **high** |

## What to do

Work through the thirteen criteria strictly in the order they are listed, and settle each one fully before starting the next. While working on one criterion, evaluate only that criterion's topic and ignore every other topic. Then assemble the report in the shape below: it leads with the findings, so a reader gets the answer on the first page, and puts the evidence after.

For each criterion:

1. Find every statement in the FACT SHEET about this topic. Quote it word for word and name the fact sheet question it answers.
2. Find every statement in the CONTRACT about this topic. Quote it word for word and give the nearest section heading or clause number above it.
3. Decide the contract commitment.
4. Decide consistency.
5. Score specificity for each document separately.
6. Note which of the criterion's warning signs actually appear.
7. Look up the severity in the table above, and show the lookup.
8. Write the finding in one or two plain sentences.

If you run out of room, finish the section you are on, stop cleanly, and tell the reader to reply "continue". When they do, resume where you stopped and carry on to the end of the report, including the closing sections.

## Rules

- Quote word for word. Never paraphrase inside a quote. If you cannot quote it, do not report it. You may join lines and collapse runs of spaces so a quote reads as one line, but no word may change, and nothing may be added or dropped.
- Silence in the contract is a real answer. Report it as absent rather than reaching for something adjacent.
- If the contract is silent, judge consistency on whether the fact sheet statements agree with each other. With one statement or none, answer aligned.
- The two documents were written at different times. When statements disagree in a way that the product changing between those times would explain, and nothing shows they describe the same moment, that is tension, not conflict — and the finding should say that age is the likely explanation. Conflict is for statements that cannot both be true of the same product at the same time, and for one part of a document taking back what another part gives.
- The warning signs in each criterion are hints, not verdicts. Report which ones appear. Their presence is evidence, not proof, and their absence does not mean the topic is handled.
- Judge the contract by what it obligates, not by what it sounds like. Language that describes, praises, or aspires is not a commitment.
- Every finding compares the two documents. Lead with what the vendor disclosed, then say whether the contract holds the vendor to it. Do not audit the contract in the abstract: a missing term matters here because a disclosure has nothing holding it up. When the fact sheet is silent too, say that neither document addresses the topic.
- Write in plain language for an agency official with no technical or legal background: short sentences, everyday words, no jargon, nothing longer than it needs to be.
- If a document says nothing about a topic, say so in that criterion's section. Do not infer, assume, or fill the gap from general knowledge about how these products usually work.
- Never assign a severity of your own. Severity comes only from the lookup table above.
- Do not draft or suggest contract language. Naming the gap is the report's job; wording to close it is a job for the government's own people.
- Weigh everything you found, but print at most the two or three most decisive quotes per document for each criterion. When statements you are not printing exist, say so and name where they sit. Never base a verdict only on what you chose to print.
- Keep each contract-commitment and consistency reason to fifteen words or fewer.

## Citing your quotes

The quote itself is the pointer: because every quote is verbatim, a reader can find it by searching the document. The locator narrows where to look.

- Locate every contract quote by the nearest section heading, clause number, or exhibit label above it, and every fact sheet quote by the question it answers.
- Do not cite page numbers. Real contracts carry several printed page numberings that disagree with each other, and different tools paginate the same file differently. If the document text contains [PAGE N] markers, ignore them too; they are extraction artifacts.
- Every quote must reproduce at least six consecutive words exactly as they appear in the document, so searching the file for it finds it.
- If a statement sits under no heading you can name, say so and describe where it falls (for example, "unlabeled paragraph following the signature block").

## The report

The report is in two parts. Part one is the answer: a summary, then every finding, most serious first, then how to validate each claim, in plain sentences a reader can act on without going further. Part two is the evidence, topic by topic, for readers who want to check. Produce it in exactly this shape:

```
# Contract review: [vendor] — [product]

[Nothing here unless something needs flagging: one short caveat if you
could see only excerpts of a document, or if the attachments were
unclear. Otherwise go straight to the summary.]

## Summary

[Two to four plain sentences: the overall verdict, leading with the most
serious findings.]

Severity counts: [N] critical, [N] high, [N] medium, [N] low, [N] none.

## Findings, most serious first

### Critical
- **[Topic]** — [the finding: one or two plain sentences comparing what
the vendor disclosed with what the contract holds it to]

### High
- [same, one bullet per criterion]

### Medium
- [same]

### Low
- [same]

### No finding
- **[Topic]** — [one clause on what the contract gets right]

[Omit any severity group with no criteria in it. Within a group, keep
the order the criteria are listed in.]

| Topic | Contract commitment | Consistency | Severity |
|-------|---------------------|-------------|----------|
| [topic] | [absent / mentioned / binding] | [aligned / tension / conflict] | [severity] |
[... one row for each of the thirteen criteria, in the order listed.
Mark "Explainability and citation" and "Independent evaluation" with an asterisk.]

* Conditional: does not always call for a contract term, and counts for less when unaddressed.

## How to validate claims

[The headings, descriptions, and topic lists below are fixed: reproduce
them as printed, keeping every topic in its group. Fill in each topic's
severity from the table above, and replace [approach] with one clause:
how to go about checking this claim, given what this review found. Name
the test, access, or commitment to seek — never draft its wording. For
topics with severity none, say what the contract already provides and
how to use it.]

### Claims you can test

A false claim shows up from the outside: run the product and watch what
it does. What the agency needs is an agreed standard and a way to test
against it.

- **Accuracy and performance** ([severity]) — [approach]
- **Robustness and failure modes** ([severity]) — [approach]
- **Bias and disparate impact** ([severity]) — [approach]
- **Explainability and citation** ([severity]) — [approach]
- **Accessibility** ([severity]) — [approach]

### Claims that need access to check

A false claim can only be caught by looking inside: a staging
environment, a configuration view, or a right to inspect. What the
agency needs is that access granted in the contract.

- **Model change and version control** ([severity]) — [approach]
- **Human oversight and override** ([severity]) — [approach]
- **Logging, audit, and public records** ([severity]) — [approach]
- **Independent evaluation** ([severity]) — [approach]

### Claims only a written promise can protect

No testing or inspection from the outside would ever reveal a false
claim. The only protection is a commitment in the contract with
consequences attached.

- **Model identity and provenance** ([severity]) — [approach]
- **Training data and data use** ([severity]) — [approach]
- **Data protection, residency, and retention** ([severity]) — [approach]
- **Liability and incident response** ([severity]) — [approach]

## Evidence, topic by topic

### [Topic] — [SEVERITY]

**What the fact sheet says**
- "[quote]" (answers: [fact sheet question])
[at most three; if more statements exist, add one line saying so and
where they sit]
[or: The fact sheet says nothing on this topic.]

**What the contract says**
- "[quote]" ([nearest section heading or clause number])
[at most three, same rule]
[or: The contract is silent on this topic.]

**Contract commitment:** [absent / mentioned / binding] — [reason, fifteen
words or fewer]
**Consistency:** [aligned / tension / conflict] — [reason, fifteen words
or fewer]
**Specificity:** fact sheet [N]/3, contract [N]/3
**Warning signs present:** [only the ones that actually appear, or "none"]
**Severity:** [contract commitment] / [consistency] / [expectation] → [severity]

[... the same section for each of the thirteen criteria, in the order
listed ...]

## Disclosure-only fields

[Which of the disclosure-only fields listed at the end of this prompt
appear in the fact sheet. These are descriptive. They are never scored
and never counted as gaps.]

## What this review cannot tell you

[The closing text at the end of this prompt, word for word.]
```

## The thirteen criteria

The fact sheet topics named below follow the published blank FactSheet template. Completed fact sheets vary: older field sets, appended questionnaires, different labels. Treat the names as topics to look for, not as labels to match exactly.

### Model identity and provenance
- Fact sheet topics: Model Information
- Assurance type: warranty. Contract expectation: required.
- What good looks like: The contract names the specific model and version family in use, names any third-party model provider and its role as a subprocessor, and requires notice before substitution.
- Warning signs: generic 'AI' or 'machine learning' with no model named; different model named in different documents; 'powered by' marketing phrasing; third-party model provider never named in the contract

### Model change and version control
- Fact sheet topics: Update procedure
- Assurance type: access. Contract expectation: required.
- What good looks like: The contract requires advance notice of material model changes, provides for re-acceptance testing after a change, and grants either version pinning or a termination right if performance degrades.
- Warning signs: updates automatically integrated; no option to revert to previous versions; updates depend on a third party's release schedule; as may be updated from time to time

### Training data and data use
- Fact sheet topics: Training Data
- Assurance type: warranty. Contract expectation: required.
- What good looks like: The contract affirmatively states that subscriber and constituent data are not used to train, fine-tune, or improve any model, with no carve-out for aggregated or anonymized use. Obligation survives termination.
- Warning signs: improve the Services; aggregated and anonymized basis; de-identified; consent granted elsewhere in the agreement; no-training promise in one section and a broad data license in another

### Accuracy and performance
- Fact sheet topics: Performance Metrics; Test Data
- Assurance type: test. Contract expectation: required.
- What good looks like: Every accuracy, uptime, or latency figure stated in the FactSheet appears in the contract as a measurable service level, with a defined test method and a remedy for failure.
- Warning signs: maintain a high level of system accuracy; based on internal testing; numeric claims present in the FactSheet but absent from the warranties; no service level agreement; no remedy or service credit

### Robustness and failure modes
- Fact sheet topics: Robustness; Poor Conditions
- Assurance type: test. Contract expectation: required.
- What good looks like: Failure modes the vendor discloses are matched by contractual guardrails: documented fallback to a human, refusal behavior for out-of-scope queries, and acceptance testing before go-live.
- Warning signs: may hallucinate; not entirely accurate; difficulty handling legally nuanced queries; disclosed limitation with no corresponding contract obligation; no acceptance testing before go-live

### Bias and disparate impact
- Fact sheet topics: Bias
- Assurance type: test. Contract expectation: required.
- What good looks like: The contract requires a documented bias assessment methodology, periodic re-testing, and grants the agency the right to conduct or commission its own evaluation using representative data.
- Warning signs: regular audits with no method, frequency, or reporting obligation; bias mitigation described only as content filtering; no baseline measurement; no agency right to evaluate

### Human oversight and override
- Fact sheet topics: Algorithmic Impact Assessment: monitoring and human override of outputs
- Assurance type: access. Contract expectation: required.
- What good looks like: The contract guarantees a human can review, override, and correct AI output before it affects a resident. Overrides are logged. No adverse action is taken solely by automated means.
- Warning signs: oversight described as feedback loop rather than decision authority; course-correct using FAQs and prompt engineering; no override logging; no due-process language for adverse determinations

### Explainability and citation
- Fact sheet topics: Explanation
- Assurance type: test. Contract expectation: conditional.
- What good looks like: The contract requires source citation on AI responses and clear labeling of AI-generated content to residents.
- Warning signs: citations provided 'when appropriate'; designed to be clear and understandable; labeling commitment present in one document only

### Data protection, residency, and retention
- Fact sheet topics: Data Protection; Jurisdiction-specific Considerations
- Assurance type: warranty. Contract expectation: required.
- What good looks like: The contract states data residency, encryption standard, retention and deletion schedule, and return or deletion on termination. Constituent data is governed by contract terms, not by an external policy the vendor can change unilaterally.
- Warning signs: privacy policy incorporated by URL; as may be updated from time to time; industry-standard security; different treatment for subscriber data and constituent data; backup disposal window inconsistent with records retention schedule

### Logging, audit, and public records
- Fact sheet topics: Ongoing Monitoring
- Assurance type: access. Contract expectation: required.
- What good looks like: The contract distinguishes security and access logs from AI decision logs (prompt, response, sources cited, timestamp), requires export in a usable format within a defined time, and preserves public-records obligations.
- Warning signs: log rights limited to access and modification logs; available upon request with no format or timeline; no retention period for decision logs; vendor confidentiality asserted over decision logs

### Independent evaluation
- Fact sheet topics: Independent Evaluation
- Assurance type: access. Contract expectation: conditional.
- What good looks like: If the FactSheet claims third-party evaluation, the contract entitles the agency to the report. If none exists, the contract requires one within a defined period or grants the agency the right to conduct it.
- Warning signs: no formal shareable studies; available upon request, respecting confidentiality agreements; willing to collaborate with third parties; customer testing offered in place of independent evaluation

### Accessibility
- Fact sheet topics: Accessibility
- Assurance type: test. Contract expectation: required.
- What good looks like: The contract names a standard and conformance level, requires a current VPAT or accessibility conformance report, sets remediation timelines, and indemnifies the agency against accessibility claims.
- Warning signs: WCAG cited without a conformance level; no VPAT or conformance report; accessibility asserted only in the FactSheet; voice or IVR accessibility unaddressed

### Liability and incident response
- Fact sheet topics: Poor Conditions
- Assurance type: warranty. Contract expectation: required.
- What good looks like: The contract allocates liability for harm caused by AI output, either by carving AI errors out of the consequential-damages exclusion or by a specific indemnity, and defines incident notification timelines.
- Warning signs: broad consequential damages exclusion with no AI carve-out; no AI-specific indemnity; disclosed hallucination risk with no corresponding liability term; no incident notification timeline

## Disclosure-only fields

Some fact sheet fields describe the product and promise nothing. Report them as disclosure only, and never count them as gaps; this is what keeps an ordinary contract from being marked mostly incomplete. They are:

- Vendor Name
- Product Name
- Overview
- Inputs and Outputs
- Optimal Conditions
- Environmental Impacts
- Responsible AI Strategy
- Purpose
- Intended Domain

## Closing text for the report

End the report with this text, word for word:

> This review checks whether the paperwork holds together, not whether the vendor told the truth. A machine reading can be wrong: check the quotes that matter before acting. This is not legal advice.

## Begin

Confirm the two attachments. Work through the thirteen criteria in the order listed, then write the report in exactly the shape above: the summary first, findings most serious first, the table, how to validate claims, then the evidence. End with the disclosure-only section and the closing text.