Skip to main content

Security and privacy

Sensitive documents deserve explicit boundaries.

A certificate of insurance carries a contractor's name, policy numbers, and carrier relationships. The design goal is that submitting one creates no lasting record of having done so.

This system is under active development and is not deployed. Each claim below is marked In place where the behaviour is implemented and tested, Built, not deployed where it is built and no environment is running it yet, or Planned where it is designed and not built. Nothing here is described as done before it is.

Your documents

Held for one request, then dropped In place

Uploaded bytes live in the memory of the single request that processes them. When a multipart upload outgrows its buffer, the web server spools it to a temporary file; every upload is closed in a cleanup block — removing that file — on success, on refusal, and on error alike. The pipeline then drops its references to the bytes and the parsed page text as an explicit step, before the response is serialized.

Documents and check results are not written to a database In place

No check history, saved contractor, uploaded document, or completed result is stored. A result exists in application memory for the length of the response, and the downloadable report is generated from it. The account database keeps only prepaid-credit accounting and the confirmed custom-rule library you choose to save, including each acknowledged requirement sentence. You can replace or delete that library from the requirements page.

Temporary storage for large uploads Built, not deployed

Uploads go straight to the backend with the request. For larger files there is a second path, written and not yet running anywhere: short-lived presigned private upload URLs into Cloudflare R2, object keys that are opaque random tokens carrying no account identifier, deletion immediately after the check, and a bucket lifecycle rule expiring anything left behind after roughly 24 hours as a backstop. Read that as designed and coded, not as protecting anything today — no bucket exists.

What is refused In place

PDF, TXT, DOCX, XLSX, PNG, JPG, and WEBP, verified by file signature or a strict text decode rather than trusting the declared type. At most 5 files per check, 25 MB per file, and 30 pages per file. Encrypted documents and documents a parser cannot read are refused. Word and Excel files are converted to bounded plain text without running macros or calculating formulas. Embedded active content is never executed. A refused file is reported as review required — never as a pass.

Third-party AI processing and retention

Field extraction is performed by a Google model API. Your documents are sent to that service, which means a third party processes them. There are two services it can be, they are the same Gemini models behind different endpoints, and their data terms differ materially — so both are described below rather than averaged into one claim. Google's Google Gemini Developer API is the default and the only one used today; a deployment can be configured to use Google Vertex AI instead for the checks it funds itself. If you supply your own key, your check always goes to the Google Gemini Developer API, because an API key is not a Google Cloud credential.

Only the documents needed for the current check are sentIn place; nothing is sent for indexing, training, fine-tuning, or any purpose beyond answering that one extraction request. A document the cheaper model could not read confidently is sent a second time to a larger Gemini model, which is a second call to the same service under the same terms, not a different one.

Documents are sent inline with the request. The provider's Files API — the upload endpoint whose stored files persist for 48 hours — is not used, so no copy of your document is created in that storeIn place.

What Google retains: the Google Gemini Developer API

This is Google's behaviour, not ours, and it is not something this system can configure away. The statements below are read from Google's own published terms, checked on 12 August 2026. They change; the authoritative text is Google's, linked at the end of this section.

  • Paid quota. Google's terms state that prompts and responses are not used to improve Google's products, and are logged for a limited period to detect and prevent policy violations, maintain safety and security, and meet legal or regulatory disclosure requirements. Google's abuse-monitoring page states that period as 55 days, covering prompts, contextual information, and output, and says that data logged for abuse monitoring is used only for policy enforcement and not to train other models. Google also states this data may be stored or cached in any country where it or its agents operate facilities.
  • Free quota. Materially different, and the difference matters for documents like these. Google's terms state that content submitted on unpaid quota, and the responses to it, are used to provide, improve, and develop Google's products, services, and machine learning technologies, and that human reviewers may read, annotate, and process API input and output — disconnected from the Google account, API key, and Cloud project first. Google's own instruction for unpaid use is not to submit sensitive, confidential, or personal information.
  • Europe. Google's terms state that for users in the European Economic Area, Switzerland, or the United Kingdom, the paid data terms apply to all services, including unpaid quota.

What this means in practice. A certificate of insurance carries a contractor's name, policy numbers, and carrier relationships, which is exactly the material Google's unpaid-quota guidance says not to submit. Platform-funded checks are therefore intended to run on paid quota, and platform-funded checks are closed until that is true of the deploymentPlanned· launch checklist. There is no way for this page to prove which quota a key belongs to, so it does not claim to: if you supply your own key, the tier of your own project decides which of the paragraphs above applies to your documents.

Two limits on the above worth stating plainly. "A limited period" in the terms is not itself a number — the 55-day figure comes from the abuse-monitoring page and is what we rely on. And neither figure is something this system verifies: it is a published commitment by a third party, restated here.

Sources, in Google's words rather than ours: Gemini API Additional Terms of Service, Abuse monitoring, and Document understanding for the Files API retention window. Content above is summarized rather than reproduced.

What Google retains: Google Vertex AI Built, not deployed

The service can be configured to call the same models through Google Vertex AI on Google Cloud. No deployment does today, and this section describes what would change if one did. Read from Google's own pages on 13 August 2026, linked at the end.

  • No unpaid quota, and so no unpaid-quota terms. This is the difference that matters most for a certificate of insurance. Every call is billed to a Google Cloud project, and the paragraph above about content on free quota being used to develop Google's products, with human reviewers able to read it, has no equivalent here. Google's Cloud terms state that it processes customer data to provide the service and does not use it to train its models without the customer's express consent.
  • Prompt logging is conditional, and the window is longer. Google states that prompts are logged for abuse monitoring only when its automated safety classifiers flag activity that needs investigation — rather than for every request — and that such logs are kept for up to 90 days, in the region or multi-region the project selected, and are not used to train or fine-tune models. So the trigger is narrower than the Google Gemini Developer API's and the retention is longer: 90 days against 55. Neither is better in every respect, which is why both are stated rather than one being called the safer option.
  • Inputs and outputs are cached by default. Google states that Gemini inputs and outputs are cached for up to 24 hours in the data centre that served the request, to speed up later prompts. That is a copy of your document's content existing for a day, and it is on by default.
  • Nothing about this is verified by us. As with the Google Gemini Developer API above, these are published third-party commitments restated here, not behaviour this system observes or can enforce.

And three things this page deliberately does not claim, because they depend on a specific Google Cloud account that this repository cannot see:

  • Whether the deploying Google Cloud account is governed by a Google Cloud Master Agreement, which Google states exempts it from abuse-monitoring prompt logging by default
  • Whether an abuse-logging opt-out has been requested and approved for the account
  • Which region or multi-region the project selects, which is where Google states any logged prompt is stored

The first of those would make the conditional prompt logging above not apply at all, which is a materially better position — and precisely the sort of favourable reading a privacy notice should not help itself to. Treat the stronger statement as the one in force unless the deployment can show otherwise.

Sources, again in Google's words: Google Cloud Platform Terms of Service, Abuse monitoring, and Generative AI and data governance for the caching window. Summarized rather than reproduced.

Using your own key Built, not deployed

You can supply your own Gemini API key, in which case the model call is billed to and governed by your own provider account and your project's tier decides its data terms. The key is accepted for one request and discarded — see API keys below. This path is built and tested in this repository; like everything else here, it is not deployed anywhere yet.

One consequence worth being explicit about: a check funded by your own key always goes to the Google Gemini Developer API, even on a deployment configured for Google Vertex AI. An API key cannot authenticate a Vertex AI request, and the alternative — running your check on the operator's cloud credentials — would bill them for a check you asked to fund. So the paragraphs about the Google Gemini Developer API are the ones that apply to a bring-your-own-key check, and your own project's tier decides which of them.

The model does not decide

This is a security property, not only a quality oneIn place. The extraction schema has no field in which a verdict could be expressed, and unknown fields are rejected rather than absorbed, so there is no channel through which a model — or text written inside an uploaded document — can express a compliance outcome.

Document content is framed to the model as untrusted evidence to be quoted, never as instructions to be followed. Snippets returned as evidence are checked against the text of the file itself for text-native PDFs, and a citation the document does not support is discarded, which unestablishes the fact that rested on it. A regression test asserts that a PDF containing injected instructions changes neither which rules run nor what they return.

API keys

The platform's model key exists only as a server-side secret, typed so that printing configuration in a log line or a traceback cannot print itIn place. It is not present in this page, in any script this site serves, or in any edge bundle. Nothing in this frontend holds a credential, because nothing in a public bundle can.

A key you supply yourself is accepted in a request header — never a URL, a query string, or a form field, all of which end up in access logs and browser history In place. It is passed as an argument for the duration of one call and is gone when that call returns; it is not stored, and it is not written to browser persistent storage. Log and exception serialization redacts key-shaped values.

What gets logged

Exactly one structured record per check, containing operational metadata and nothing read out of a document In place. Running a service responsibly needs latency and error rates; it does not need your contractor's policy number.

Recorded

  • Request identifier
  • Whether the platform key or your own key was used
  • Which model API served the check
  • Number of documents and total page count
  • Model name and number of model calls
  • Whether a document had to be read again by a larger model, and why, as a code
  • Latency and token counts
  • Counts of passed, failed, and review results
  • An error code, when something failed

Never recorded

  • Document contents, in whole or in part
  • Extracted names, policy numbers, limits, or dates
  • Request or response bodies
  • Anything shaped like an API key

Identity and abuse controls

Signing in exists for one reason: to attach a small allowance of platform-funded checks to a durable identity. There are no roles, organizations, or profiles. The design uses Auth0 for the sign-in itself, so no password reaches this system, and stores only a counter — keyed by the account identifier — of checks used, alongside bot verification and rate limits at the edgeBuilt, not deployed. That counter is not a profile and not an activity log: it records how many checks have been run, not what was in them. Currently 3 platform-funded checks per account is the intended allowance.

This website

Google Analytics (gtag.js) provides aggregate pageview telemetry. No advertising trackers, no cross-site pixels, and no externally hosted fonts. Interactive client-side code exists only where workflow interaction genuinely requires it, such as the upload wizard.

What this does not claim

  • It is not a verification service. A check reads the documents you supply; it does not confirm with a carrier that a policy is in force, and it does not detect a convincing forgery.
  • It is not legal, insurance, or compliance advice, and a compliant result is not a certification.
  • It is not a substitute for human review. Where the documents do not establish a fact, the result says so and asks a person to look.
  • It has not been through an external security audit or penetration test.