AI vendor security questionnaire for SaaS and agent tools

A practical questionnaire covering AI data flows, model providers, connectors, agent actions, retention, evaluation, monitoring, incidents, and exit.

Security review mapping an AI vendor's models, data, connectors, actions, and logs
Sources4 other
Decision next step

Compare the tools behind this article on ToolVerse.

Open ToolVerse for evidence, pricing context, alternatives, and current review status. Every link below navigates to the external ToolVerse directory.

Compare AgentShield and Promptfoo and Guardrails AI Open on ToolVerse · external

Quick answer

A conventional SaaS questionnaire may miss model providers, prompt logs, retrieval indexes, tool calls, and autonomous actions. Use the questions below to expose those paths and request artifacts—not marketing assurances. Security and procurement teams should scale the depth of review to data sensitivity, action authority, and failure impact.

Architecture and data

  1. Diagram every service that receives prompts, files, retrieved content, outputs, traces, or feedback.
  2. Identify model providers, regions, subprocessors, and tenant-isolation boundaries.
  3. State whether any customer data trains or improves models by default or opt-in.
  4. Provide configurable and maximum retention for each data copy, including backups.
  5. Demonstrate deletion of source data, embeddings, caches, logs, and derived examples.

Identity, connectors, and actions

  1. List connector scopes and whether read, write, delete, and admin access are separable.
  2. Explain how tokens and secrets are stored, rotated, redacted, and prevented from entering prompts.
  3. Show approval, transaction-limit, allowlist, dry-run, and rollback controls.
  4. Describe defenses against prompt injection, confused-deputy behavior, and cross-tenant retrieval.
  5. Provide an audit trail linking user intent, model decision, tool input, action, and result.

Quality and change control

QuestionStrong evidenceWarning sign
How is quality evaluated?Versioned datasets, rubrics, critical gatesHand-picked demos
What happens on model change?Notice, regression tests, rollbackSilent replacement
How are unsafe outputs handled?Layered controls and measured failure casesA generic safety claim
How are incidents learned from?Postmortems, corrective actions, customer noticeNo AI-specific process

Example: no training is incomplete

A vendor states that customer data is not used to train models. The product still retains prompts for support, copies traces to an observability provider, and stores embeddings indefinitely. The answer may be technically true but insufficient. Ask about every use and copy, then align it with the retention policy guide.

Evidence to request

Request current architecture and data-flow diagrams, security audit scope, penetration-test summary, subprocessor list, incident policy, business-continuity evidence, deletion procedure, model-change policy, and a completed action-control matrix. For high-risk agents, observe a scoped test rather than accepting screenshots. After selecting the questions, use a governed security questionnaire automation workflow to keep evidence owners, approvals, freshness, and exceptions visible as answers are maintained.

Compare relevant controls and evaluation tools in AI security questionnaire tools. AgentShield and Promptfoo address different parts of testing; neither is a substitute for vendor evidence.

Risk-tier the questionnaire

An assistant that drafts public copy needs a different review from an agent that can read customer records and issue refunds. Classify sensitivity, action authority, user population, scale, reversibility, and regulatory impact. Low-risk tools may use a concise questionnaire; high-risk agents need architecture review, technical testing, contract commitments, and a staged production exercise.

Map each answer to the exact feature and plan being purchased. Vendors often describe an enterprise control that is not available in the proposed tier, or a model-provider policy that changes when optional features are enabled. Ask for configuration screenshots or a demonstration where evidence is inexpensive, and place durable commitments in the agreement.

Agent-specific scenarios to test

Give the system retrieved content containing instructions to disclose data or call a tool. Confirm that untrusted content cannot override the user’s approved intent. Test a user with insufficient permissions, a revoked connector token, an ambiguous destructive action, a tool timeout after a possible write, and a request that exceeds a transaction limit.

Observe the trace. It should show identity, authorization, model and policy decision, tool input, external result, approval, and final state without exposing secrets. Confirm operators can distinguish a failed action from an unknown outcome and can reconcile or roll back safely. Ask how the vendor validates its own tool descriptions and connector updates.

Operational security

Review secure development, dependency management, vulnerability disclosure, penetration testing, tenant isolation, encryption, key management, employee access, and disaster recovery in the context of the AI architecture. Model gateways, vector stores, browser infrastructure, and observability services may expand the traditional application boundary.

Ask who monitors model abuse, prompt injection, anomalous tool use, data exfiltration, and repeated policy denials. Define notification timing and content for incidents involving prompts, model providers, connectors, or agent actions. A generic breach notice may not cover an unsafe external action with no loss of stored data.

Review and renewal

Track unknowns, exceptions, compensating controls, owners, and expiry dates. Before rollout, confirm the deployed configuration matches the reviewed design. Reopen the assessment for new models, subprocessors, connectors, action scopes, retention changes, or acquisitions. At renewal, request updated evidence and compare real incident, usage, and evaluation results with the original claims.

The questionnaire is a decision record, not a score-generating ceremony. A vendor can be acceptable with transparent limitations and strong compensating controls, while a vendor with many polished answers can remain unsuitable if a critical data path is unverifiable.

Grade the assurance, not the prose

Use five states: verified by evidence, contractually committed, vendor-attested, planned, and unknown. “Yes” is not a state. A control may be not applicable, but the reviewer should explain why the architecture does not require it. This vocabulary prevents polished questionnaire language from flattening differences in assurance.

For each high-risk unknown, block, test, add a compensating control, negotiate, or accept with an owner and expiry. Disabling write scopes, excluding sensitive data, limiting users, shortening the pilot, or retaining manual approval can bound risk. Provide the vendor an exact architecture and use-case description so answers apply to the real deployment rather than a hypothetical generic product.

Ask the vendor to correct the completed record when architecture or terms change. Version the questionnaire and retain the evidence date; security pages are living documents, not proof of what applied at purchase. An annual refresh is a minimum, while new action scopes, model providers, or sensitive data should trigger review immediately.

Decision

Classify each answer as verified, contractually committed, planned, unknown, or not applicable. Block production when a material data path or action boundary remains unknown, and carry accepted exceptions into the contract and risk register.

Build the shortlist

Compare the referenced tools side by side.

Compare AgentShield and Promptfoo and Guardrails AI →