Most evaluations of an AI system stall before the first question is asked. The team that wants to test the system needs documents to test it on, and the documents that matter are the ones an organization is least able to hand over. Releasing them to a vendor, or to an internal pilot team, requires a data-sharing agreement, a security review of the test environment, and a decision by someone with the authority to accept the risk. Those steps often take longer than the evaluation itself, and until they are complete nobody learns whether the system can do the work.
A synthetic test set removes that dependency for most of the evaluation. It is a collection of invented documents, built to resemble the real ones in structure, volume, and difficulty, together with a list of questions whose correct answers are known in advance. Because no real record is involved, the set can be shared without approval, loaded into any candidate system, and reused for as long as the workflow exists.
Contents of a Useful Test Set
A test set is useful to the degree that it reproduces the conditions that make the real work difficult. Every party in it is fictional, which means each company, person, address, and account number is invented and the names are checked against public registries so that none belongs to a real entity. The documents follow the structure of the real collection. If that collection holds scanned amendments, multi-page tables, and email threads, the synthetic one holds the same types in similar proportions, and it is large enough that the system has to retrieve the right passage from among many plausible ones.
The set also contains hard cases that were placed there on purpose. Real collections include records that contradict each other, near-duplicate versions of one document, fields that were never filled in, and questions from users that the collection cannot answer. A test that leaves these out measures only the easy part of the work.
Finally, every test question has a recorded correct answer. The record gives the answer, the document and section that support it, and the behavior expected of the system, which may be to answer, to report a conflict between sources, or to decline.
A Worked Example With Supplier Contracts
Take a fictional manufacturer of industrial pumps that wants an assistant for questions about its supplier contracts. The manufacturer is given no real name, and its suppliers are identified by invented names and supplier numbers. The specification for the test set describes a company with forty suppliers. Each supplier has a master agreement, between one and five amendments, and several years of purchase orders and quality notices. Generators produce about 1,200 documents from that specification, varying the layout, length, and wording so that no two agreements look alike.
Several hard cases are written into the specification. One supplier has two amendments that state different payment terms, with effective dates three months apart, so the correct answer depends on the date a question refers to. Another has a draft and a signed version of the same agreement that differ in a single clause on late-delivery charges. A group of purchase orders has no delivery date. The question list includes requests about a supplier that does not appear in the collection and about subjects the collection does not cover, such as employee pay.
The answer file holds 150 questions. The question about payment terms is marked as requiring the later amendment, with a statement that it replaced the earlier one. The question about the absent supplier is marked as one the system should decline to answer.
Building the Set
The work produces four artifacts, and the order matters. The first is a written specification of the fictional world, covering the entities, their relationships, the time span, the document types, and the hard cases. It should be drafted with the people who do the real work, because they know which situations cause mistakes.
The second is a set of generators, usually templates combined with a language model, that produce documents from the specification. They draw on the specification alone and never on real records. The third is the ground-truth file, which is written from the specification and not from the generated documents, so that the correct answers do not depend on any model’s reading of the text. The fourth is a scripted question set that runs unattended against any system and is scored the same way each time.
Measurements
Seven measurements cover most of what a buyer needs to know. Answer accuracy is scored against the ground-truth file. Citation correctness checks whether the passage a system cites is the one that supports its answer. Refusal is scored on the out-of-scope questions, where a confident answer counts as a failure.
Behavior at volume is tested by running the set at its original size and again after the generators expand it, for example from 1,200 documents to 50,000, because retrieval that works on a small collection can degrade on a large one. Response time and cost per answer are recorded for every run. When more than one model is under consideration, each is run on the same set, which makes differences between models visible on identical work.
Results are most useful when they are reported by question type. An average can conceal consistent failure on conflicting records while routine lookups score well.
Running the Test on a Governed Platform
A test is more informative when it runs under the controls the production system will have. That calls for an isolated environment with the same access rules, the same guardrails, and the same audit logging as production. The evaluation then shows whether a user in one role can retrieve a document restricted to another, and whether the log records what each question retrieved.
Knowledge Spaces, Sprinklenet’s platform, is offered as a hosted service and provides role-based access, configurable guardrails, scored evaluation runs, and audit logging. A synthetic set can be loaded into a separate space on the platform and tested with no connection to production data.
Limits of Synthetic Testing
Synthetic data does not prove performance on real records. Real documents carry scanning defects, local abbreviations, and drafting habits that a specification will not fully anticipate. A short confirmation run on real data, under the approvals that real data requires, is still necessary before production use. The synthetic phase makes that run shorter, because most problems have already been found and fixed.
The same caution applies to how the set is constructed. It must not be built by copying real records and changing the names, since the structure and detail of a real record can identify its subject after the names are gone.
Regression Testing After Launch
The test set remains useful once the system is in service. Models are retired and replaced, prompts are revised, and retrieval settings change, and any of these changes can alter answers that were previously correct. Rerunning the scripted question set after each change shows, before release, whether accuracy, citation, or refusal behavior has moved and on which question types. Failures observed in production can be added to the specification as new hard cases, so the set becomes more demanding as the system matures.
Discuss an evaluation on synthetic data with Sprinklenet.
Related reading: RAG Evaluation: What to Measure Before Launch; How to Build a Governed AI Sandbox; AI Governance and Policy Services.
Founder and CEO, Sprinklenet
Jamie Thompson is founder and CEO of Sprinklenet, where he leads AI implementation, systems integration, and Knowledge Spaces delivery for regulated and operational teams.
His work focuses on moving AI from strategy and pilot activity into governed production systems with clearer retrieval, workflow, evaluation, and audit controls. LinkedIn profile.

