The pilot graveyard is an evaluation problem
Most stalled AI projects never defined what "correct" looks like. A 300-example golden set, built with the people who do the job, is the cheapest investment in your AI programme.
We design, build and operate agentic AI systems for banks, insurers, hospitals, universities and manufacturers — with the evaluation, observability and human approval gates that production demands.
Six capabilities, delivered as working systems rather than slide decks. Each one ships with evaluation sets, monitoring and a clear operating model.
Multi-step agents that plan, call tools, read your systems and complete tasks end-to-end — with approval gates where the risk warrants it.
Turn forms, contracts, claims packs, lab reports and invoices into reviewed, structured data with field-level confidence and audit trails.
Grounded answers over policies, manuals and case history — hybrid retrieval, re-ranking and citations so every answer can be traced to a source.
Offline eval sets, online monitoring, prompt and model versioning, drift and cost dashboards — the plumbing that keeps AI reliable after launch.
Fine-tuned small language models and on-premise deployments for teams that cannot send data out, or need predictable unit economics at scale.
Model inventories, risk tiering, decision logs and board-level reporting aligned to MAS AI risk guidelines, PDPA, the India DPDP Act and the EU AI Act.
Two of five anonymised engagements from the last 18 months.
Motor and property claims arrived as mixed PDFs, photos and emails. Assessors spent most of their day re-keying data, and complaint volumes were rising with the backlog.
A claims intake agent that classifies the pack, extracts fields with confidence scores, retrieves the policy terms and drafts an assessment. Anything below threshold, or above a payout limit, routes to a human with the full trace attached.
Corporate onboarding needed documents from six jurisdictions reviewed against internal policy. Reviewers were inconsistent, and turnaround regularly breached the committed SLA.
A document-review copilot grounded on the bank's own policy manuals, with a checklist agent that drafts the review memo, cites every clause it relied on, and flags gaps for the analyst to confirm.
A fixed-scope path from a real business problem to a system your operations team runs. No open-ended discovery phases.
We map the workflow, the data, the risk tier and the economics. You get a build/no-build decision with a business case.
A working agent on your data with an evaluation set, approval gates and a measured baseline — not a chatbot demo.
Integration with your systems, hardening, observability, runbooks and training for the team that will own it.
Managed service or handover. Weekly eval review, drift and cost monitoring, model upgrades without regressions.
Short positions from our current engagements.
Most stalled AI projects never defined what "correct" looks like. A 300-example golden set, built with the people who do the job, is the cheapest investment in your AI programme.
For narrow, high-volume tasks, a fine-tuned 7–14B model on your own infrastructure now beats frontier APIs on cost, latency and data-residency — if you have the evaluation discipline to prove it.
MAS guidance, the EU AI Act and India's DPDP rules converge on one practical demand: show the decision, the data it used and who approved it. Build the log before you build the agent.
Bring one workflow that costs you time or money. In 45 minutes we'll tell you whether an agent can take it on, what it would need, and what it would cost — honestly.