SOC 2 for AI Startups

SOC 2 for AI companies, model-aware.

AI startups have audit exposure traditional firms don't understand: foundation-model sub-processors, training-data controls, prompt logging, and vector-store isolation. Skadi is built for it. Our senior CPAs scope around model architecture, and we audit with a privacy-first local model so your evidence never leaves for a third-party LLM.

4.9on Trustpilot
Independent SOC 2 Type II
Why this vertical is different

Audit challenges we solve for you.

Foundation model sub-processors

Every hosted LLM you call is a sub-service org. Your system description, security review, and monitoring evidence must reflect that — and buyers now ask specifically about enterprise-tier data terms.

Training & fine-tuning data

Customer data used for embeddings, RAG, or fine-tuning creates retention, lineage, isolation, and deletion obligations most SOC 2 templates miss entirely.

Evidence privacy paradox

You sell privacy to customers, then hand screenshots and logs to auditors using public LLMs. Skadi's local-model audit closes that gap by design.

Fast-moving surface

Prompts, model versions, and eval suites change weekly. Change-management controls that were fine for a stable app fail here without adaptation.

Where auditors focus

Controls that matter most for your business.

Model & inference access (CC6.1)

Who can call production models, deploy new prompts, or change fine-tunes. Role-based access, approvals, and audit trails on model gateways.

Training-data governance (CC6.7, C1.1)

Data lineage, classification, opt-out handling, and deletion evidence — especially for tenants who contractually forbid their data being used for training.

Tenant isolation on vector stores (CC6.6)

Namespace or per-tenant index partitioning, prompt-context isolation, and evidence one customer's embeddings cannot surface in another's session.

Prompt/response logging & retention (CC7.2, C1.2)

Policy, technical enforcement, and retention windows aligned to customer contracts; access controls on logs that may contain sensitive inputs; scheduled deletion evidenced.

Sub-processor management (CC9.2)

Formal inventory of OpenAI, Anthropic, Pinecone, Weaviate, and every inference/vector vendor — with DPA, enterprise-tier data terms, security review, and monitoring evidence.

Model change management (CC8.1)

Approvals and rollback for prompt changes, model version bumps, and eval-gated releases — the fastest-changing surface in an AI company, and the one auditors most often find under-controlled.

Model evaluation & safety gates (CC7.1)

Documented evaluation cadence for safety, bias, and jailbreak resistance where contractually promised, with release gating evidenced from CI.

Data deletion in derived artifacts (C1.2, P4.2)

Deletion propagated to embeddings, fine-tune datasets, cached prompts, and evaluation corpora — not just the source database.

Common Type II findings

What actually fails in this vertical.

Real deviations we see in SOC 2 Type II fieldwork — and what your platform and security teams should fix before the observation window starts.

Foundation model on default (non-enterprise) tier

Production traffic routed through OpenAI/Anthropic default terms where inputs may be retained. Enterprise/zero-retention tier not enabled — the single most-cited AI-specific finding in 2025.

Evidence to collect
  • Executed enterprise/zero-retention agreement with each model vendor
  • Screenshot of API org settings showing zero-retention flag enabled
  • Egress config or model-gateway policy pinning production to the enterprise endpoint
  • Sub-processor register listing each model vendor + data-handling tier
Vector index shared across tenants without partition key

Single Pinecone/Weaviate index storing embeddings for multiple customers with only application-layer filtering. One query bug and you have a cross-tenant leak — auditor calls this a design deficiency.

Evidence to collect
  • Architecture diagram showing per-tenant namespace / index
  • Terraform or vector-DB config snippet enforcing tenant filter
  • Automated isolation test in CI with passing run logs
  • Threat model doc covering cross-tenant retrieval risk
Prompt logs retained past customer-committed window

Contract promises 30-day deletion; logs sit in an S3 bucket for a year because no lifecycle rule was set. Deletion must be technical, scheduled, and evidenced.

Evidence to collect
  • S3 lifecycle policy / log-store retention config
  • Retention policy document referencing customer commitments
  • Sample deletion-job log for the window
  • MSA/DPA excerpt with the committed retention window
Fine-tune datasets include opted-out customer data

Opt-out flag exists in the app DB but the training pipeline reads a raw dump that ignores it. Data-lineage control fails on first walkthrough.

Evidence to collect
  • Data-lineage diagram from source to training corpus
  • Pipeline code / DAG showing opt-out filter
  • Sample training-dataset manifest with opt-out counts
  • Ticket for opt-out reconciliation run during the window
Prompt/model changes deployed without approval trail

Prompts edited directly in a config dashboard, no PR, no reviewer, no rollback plan. Change management (CC8.1) exception on nearly every AI startup's first audit.

Evidence to collect
  • Git history for prompts/model configs under version control
  • PR with reviewer approval + CI evals passing
  • Change-management policy covering prompts and model versions
  • Rollback runbook with a recent execution log
No evidence of eval gates before model release

Safety and regression evals exist but running them is not enforced in CI. Marketing claim exists ('we test for jailbreaks'); control to back it does not.

Evidence to collect
  • CI workflow file with required eval job (branch protection references it)
  • Eval report artifact for a sampled release
  • Written eval policy with pass/fail thresholds
  • Release ticket linking to eval artifact and approver

Why traditional auditors miss the risk

Most audit firms wrote their SOC 2 playbooks before generative AI existed. They treat an inference API as 'just another SaaS vendor' and miss the actual risks: training-data leakage, prompt injection into logging pipelines, and customer contracts that forbid third-party model training. Skadi's methodology was rebuilt around modern AI architecture from the ground up.

The AI-company evidence problem

AI companies face a specific paradox at audit time: they sell privacy and data protection, then are asked to upload screenshots and system exports to consultants who process them with public LLMs. Skadi's audit is performed by a local model running inside the audit environment. Evidence stays put, and the CPA reviews it in place — the same posture your customers demand of you.

Enterprise readiness for AI buyers

Enterprise procurement is now asking AI vendors questions no one asked two years ago: model provenance, training opt-outs, prompt logging retention, tenant isolation on vector stores, and deletion in embeddings. A SOC 2 Type II report that speaks to those specifically — instead of pretending the AI layer doesn't exist — is a practical way to unlock those deals.

Preparing for ISO 42001 and AI-specific frameworks

ISO/IEC 42001 (AI management systems) and the NIST AI RMF are becoming the next enterprise ask. We scope SOC 2 so the same governance, risk-assessment, and monitoring evidence extends cleanly to a 42001 program later. One control library, multiple attestations — instead of duplicated work across frameworks.

FAQ

Questions we hear from teams like yours.

Ready to scope your SOC 2 Type II?

Book a call and we'll walk you through timeline, scope, and pricing for your company.

Back to home