← All case studies
Generative AI
A hallucination-detection framework for GenAI
A team validated a hallucination metric-and-diagnosis framework across multiple models to detect and measure incorrect LLM outputs before production.
Reusable
safety layer for GenAI
The challenge
Scaling GenAI into critical workflows risked embedding errors. Existing measures were immature, detection leaned heavily on manual review, and early experiments were inconclusive.
How it ran on NayaOne
1
Controlled environments
Multiple secure sandboxes ran experiments across different models without exposing data or production systems.
2
Metric validation
Context relevance, faithfulness, answer relevancy and summarisation metrics tested against human evaluation.
3
Calibration & stress
Model-specific calibration explored to improve accuracy and reduce reliance on human oversight.
Outcomes
Multi-model
evaluation on identical query-context-response triplets
Human-aligned
metrics benchmarked against reviewers
Reusable
safety layer for future deployments