Claims – Synthetic Medical Reports
1,237 fabricated claim documents across 30 case packages – Auto, Workers’ Comp and Liability.
Synthetic Data
Evaluations stall on data, not technology. NayaOne gives your teams a library of synthetic and public datasets – bank statements, identity documents, claim files, fraud, loans, on-chain activity – spanning 21 ecosystems and seven regions, loadable into a sandbox in minutes.
200+
synthetic and public datasets
21
ecosystems, from payments to ESG
7
regions covered worldwide
0
real customer records exposed
ON THIS PAGE
The blocker
Data library
Coverage
The principle
A dataset in depth
FAQ
THE BLOCKER
Nobody loses six months to the technology. They lose it to the data.
By the time a vendor is shortlisted, the decision is usually already made in principle. What follows is the wait: an internal data request, a privacy review, a redaction exercise, a sign-off. The vendor demos on its own clean sample instead, and nobody learns anything.
01
Access takes longer than the test
Production data needs a legal basis, an environment and an owner. Each one adds weeks before a single result exists.
02
Vendor demos run the happy path
Every tool works on the sample it chose. You find the failure modes in month four of production, not in the evaluation.
03
Comparisons aren’t like for like
Three vendors, three datasets, three sets of numbers. Nothing in the pack a committee can actually decide on.
The data library
A selection of the library as it appears inside the platform. Filter it the way your teams do, then load any dataset into a sandbox workspace.
1,237 fabricated claim documents across 30 case packages – Auto, Workers’ Comp and Liability.
Statistically representative businesses and their accounts-receivable ledgers.
Synthetic agents responding to inbound customer contacts across channels.
Work schedules and shift patterns for synthetic customer support agents.
Skill matrices for each synthetic support agent, mapped to queue routing.
Completed training records for synthetic support agents.
Aggregate FCA complaints data, 2021 H2, by firm and product category.
DB1B origin-and-destination survey: a 10% sample of reported airline tickets.
PM2.5 mean annual exposure, micrograms per cubic metre, by country.
The 44 datasets above are a representative slice for browsing. Inside the platform your teams see the full library of 200+ synthetic and public datasets across the same 21 ecosystems and seven regions, and bespoke sets are generated to a specific use case in about a week.
This is a sample, not the catalogue.
The datasets above are a representative slice for browsing. Inside the platform your teams see the full library of 200+ synthetic and public datasets across the same 21 ecosystems and seven regions, and bespoke sets are generated to a specific use case in about a week.
See the full library →
WORLDWIDE COVERAGE
A UK payment rail behaves nothing like a Gulf one. The data reflects that.
Regulation, payment infrastructure, identity documents and customer behaviour differ by market. Global institutions evaluate once and deploy in many, so the library carries region-specific sets – ISO 20022 in Europe, UPI in India, CDR consents in Australia, mobile-money agents in Africa – rather than one anglocentric sample stretched to fit.
Europe
19 datasets · 18 ecosystems
Synthetic UK Business Accounts Receivables · Customer Support Agent Database · Synthetic Agent’s Schedules
North America
11 datasets · 14 ecosystems
Claims – Synthetic Medical Reports · Global Airline Flight Itineraries · E-Commerce Toy Listings on Amazon
Middle East
3 datasets · 5 ecosystems
Global AML Risk Scores · GCC Identity Documents · Tokenised Fund Registry
Asia
4 datasets · 8 ecosystems
UPI Transaction Corpus (India) · APAC Private Banking Client Book · Trade Finance Document Bundles
Oceania
3 datasets · 6 ecosystems
Motor & Home Policy Book · Australia CDR Consent Records · SMB Cashflow & Invoicing
South America
2 datasets · 5 ecosystems
Brazil PIX & Consumer Credit · Merchant Onboarding Files
Africa
2 datasets · 3 ecosystems
Global Air Pollution Exposure · Africa Mobile Money Agents
Talk to data engineering →
Your market missing?
We generate to jurisdiction, language and regulatory shape.
Talk to data engineering →
The principle
A vendor demo runs on the happy path. Real production doesn’t. The point of synthetic data isn’t a tidy spreadsheet, it’s representative data with the missing records, fragmented histories and multi-provider gaps your teams actually face, so you find out how a tool behaves against reality before you commit.
Edge cases on purpose. Missing fields, incomplete records, anomalies and conflicting sources, the conditions that break naive tools.
Real coverage. Built around genuine personas, scenarios and geographies, so “representative” actually means something.
Straight into beta. Realistic personas remove the single biggest blocker to getting a working build in front of pilot users.
synthetic_dataset · caregiver_v3
VALIDATING
patient_0142
patient_0145
patient_0148
patient_0151
patient_0154
patient_0157
patient_0160
Populated
Intentional gap
Anomaly
What our customers say
“Our collaboration with NayaOne has dramatically streamlined how we vet fintech vendors, positioning us well ahead in the digital transformation and AI race.”
Head of Innovation
Top-10 global bank
“Test data is a challenge all the time, and we’ve got to get systems in sync. This removes that friction at the exact moment we need to move.”
CIO, Group Insurance
North American insurer
“NayaOne shows what’s possible when ambition meets execution, turning bold ideas into measurable industry impact.”
Innovation Director
European financial group
In practice
A North American insurer building a new caregiver proposition needed to test an AI assistant against the reality its users face, not a clean demo set. NayaOne generated a synthetic healthcare dataset spanning 25 caregiver personas across multiple health systems and states.
It deliberately included missing records, fragmented histories and multi-provider gaps, so the team could prove the AI held up against messy real-world conditions, then move straight into beta with realistic personas. No real patient record was ever exposed.
1 week
to generate the synthetic dataset
3 days
to a ready environment with the tech stack
25
caregiver personas across real scenarios
0
real patient records exposed
A dataset in depth
Claims – Synthetic Medical Reports: what a purpose-built evaluation dataset looks like.
A North American insurer set out to test AI-assisted claims summarisation. Adjusters were manually reviewing large, complex bundles of medical records and bills across workers’ compensation, auto and liability – slow work, and easy to miss a billing code that doesn’t match the reported injury.
Real patient records could not leave the estate, so the evaluation needed 100% synthetic data. NayaOne built a complete claim corpus: every vendor tested on the same 30 cases, the same documents, the same deliberately awkward edge cases.
1,237
documents in the corpus
30
complete claim packages
3
claim lines: auto, WC, liability
14+
document types modelled
18/18
automated quality checks passing
Fourteen document types, including handwriting
Each case folder mirrors what a claim handler actually receives – official billing templates alongside handwritten legal-pad notes, specifically to test extraction under realistic conditions.
HCFA-1500
~220 · official template, all 33 boxes
UB-04
~50 · AHA facility billing form
PCP SOAP notes
Handwritten · pen jitter to challenge OCR
Chiropractic SOAP
Handwritten · per-visit codes
Radiology reports
ACR dictation format
Operative reports
OR dictation, findings, complications
PT / OT sessions
Functional outcome scores per session
EMS / PCR
NEMSIS-style pre-hospital reports
IME reports
Causation, apportionment, MMI
Work status / RTW
Restrictions and effective dates
Neuropsych evaluation
Standardised scores, TBI causation
Discharge, consults, DME
Multi-specialty letters and invoices
A deliberate difficulty spectrum
Cases run simple to complex on purpose: bilateral surgery, disputed causation, traumatic brain injury with neuropsychological testing.
Simple
5 cases
Moderate
12 cases
Complex
13 cases
See a real case package
WC_01 – lumbar strain, electrician
Workers’ compensation · 27 documents · simple complexity. One NayaOne-branded PDF combining the case summary, mechanism of injury, provider notes and the full billing trail.
Open the sample scenario (PDF) →
Entirely fabricated. No real patient, provider or claim data was used at any point.
Why it worked as an evaluation.
Every vendor received identical inputs, so scores could be compared directly – extraction accuracy on handwriting, whether mismatched billing codes were caught, and how each handled the deliberately awkward cases.
Q.
Is synthetic data actually representative?
It replicates the statistical properties and edge cases of real datasets, distributions, correlations, anomalies, so models and integrations behave as they would in production, without the exposure.
Q.
Can we bring our own data instead?
Yes, under your controls, inside the air-gap. Most teams start with synthetic to move fast, then introduce their own data once the approach is proven.
Q.
Does this satisfy our privacy and compliance teams?
Synthetic data carries no real customer information, so validation happens without privacy risk. It’s why regulators, including the FCA Digital Sandbox, rely on the approach.
Q.
What if the dataset we need doesn’t exist yet?
Most substantial evaluations end up with a bespoke set. Data engineering builds to your scenario, jurisdiction and document formats, typically in about a week, and it stays available for every subsequent evaluation you run.
Q.
Can several vendors be tested on the same data?
That is the point of it. Every vendor receives identical inputs in identical environments, so the scores are comparable and the evidence pack stands up to a committee.