Strategic Sourcing GuideDecisions. Evidence. Leverage.
AI & sourcing

Buy the AI capability. Verify the operating boundary.

Evaluate the exact product, model dependencies, data behavior and cost per useful task—not the general promise of AI.

Strategic Sourcing GuideReviewed 4 October 2026Independent buyer guidance
Buyer principle

Separate what is demonstrated, contractually committed and still dependent on a provider policy.

The AI supply and responsibility chain

17 / Map the AI responsibility chain
01Your workflowData, users, permissions
02ApplicationOrchestration and storage
03Model & servicesInference, retrieval, tools
04OperationReview, monitoring, change

An AI application may depend on model providers, cloud infrastructure, retrieval services, connectors and human review. Map who processes data and who controls each change. A product subscription, API integration and self-hosted model have different operational and contractual boundaries.

An evidence agenda for AI procurement

AI sourcing review
AreaEvidence to requestCommercial consequence
Data useProduct-specific terms for prompts, outputs, training, retention and deletion.A general no-training statement does not settle storage or feature-specific processing.
Model changesVersioning, retirement notices and migration responsibilities.A model change may require revalidation and integration work.
QualityResults on buyer-defined tasks, error analysis and refusal behavior.Benchmarks do not establish performance on your workflow.
Security / toolsPermission boundaries, connector scope and adversarial testing.An agent with action permissions can create consequences beyond a text error.
EconomicsMetering, retries, supporting services and review effort.Token price is only one component of cost per useful outcome.
ExitExportable data, prompts, evaluations and service dependencies.A replaceable model does not necessarily make the application portable.

Use the voluntary NIST AI RMF and its Generative AI Profile as risk-management references, not as product certifications. Bring security, privacy, legal and operational owners into the review for the actual use case.

Run a bounded, decision-oriented pilot

Define the tasks, representative inputs, prohibited data, success thresholds and human review. Include failures and difficult cases, not only curated examples. Measure accuracy appropriate to the task, review time, latency, completion rate and total consumption. Test whether the system signals uncertainty usefully.

For tools that can act, use a controlled environment and least-privilege permissions. Decide which actions require human approval. A successful summary task is not evidence that the same system can safely execute purchases or change supplier records.

Current policy checks

Reviewed 4 October 2026. Provider policies must be verified at purchase and again when a product or feature changes. Anthropic publishes a model deprecation policy and a separate explanation of zero-data-retention scope. These illustrate two distinct diligence questions: how long a capability remains available, and exactly which processing is covered by a data arrangement. Do not generalize either policy to another provider or product.

Ask the supplier to identify the applicable documents and any feature, model or deployment exceptions in writing. Capture the versions relied upon and obtain qualified review where terms are material. Emerging agent capabilities and vendor roadmaps should be evaluated separately from delivered, measured functionality.

Set a pilot decision rule before seeing results

For a proposal-extraction assistant, build a test set containing ordinary responses, conflicting terms, scanned tables and deliberately unanswered requirements. Have a qualified reviewer establish the expected extraction. Measure missing facts, incorrect facts, unsupported citations and the time needed to correct the output. A summary that sounds convincing is not the success criterion.

Set acceptable thresholds based on the consequence of the task. A draft internal summary can tolerate a different review burden from a workflow that updates supplier records. Record failures by type so that a good overall average does not hide the category of error that matters most.

  • Require an explicit unknown result when the source does not answer the question.
  • Verify whether citations identify the actual supporting passage.
  • Check the cost of retries and human corrections, not just the first request.
  • Repeat critical cases after a material model, prompt, connector or policy change.
  • Keep an exit path for the pilot and prevent trial access from becoming an unmanaged production service.

A supplier’s benchmark can inform the test design, but the award should rely on evidence from the intended workflow and the agreed operating controls. If the system cannot meet the bounded task reliably enough, narrow the scope, increase review or decline that use case.

Sources & context

The decision frameworks and illustrative examples are original editorial guidance. Sources support the stated context; they do not endorse this guide.

Continue the work