Skip to content
RegulensR
AI & automation

Grounding is the whole game: why retrieval quality decides whether AI answers about Indian law are safe

The difference between a useful AI assistant for regulatory questions and a dangerous one is not the underlying model. It is whether the answer is built from retrieved source text or from the model’s general training — and for Indian state-level law specifically, that difference is larger than most evaluations account for.

Rohit MenonPrincipal Regulatory Analyst5 min read0 views

Ask a general-purpose AI assistant, with no specific regulatory grounding, a question about Indian compliance obligations, and it will very likely answer fluently, confidently, and — with a frequency that should concern anyone relying on it — incorrectly. This is not a criticism of any particular model's capability. It is a structural consequence of how these models are trained and what they are being asked to do, and understanding the mechanism is what separates a genuinely useful regulatory AI tool from a dangerous one wearing a helpful interface.

Why general training is the wrong foundation for this specific task

Large language models are trained on broad text corpora that are not evenly distributed across jurisdictions, and are heavily weighted toward the volume of English-language legal and regulatory text available online — which skews toward US and, to a lesser extent, UK and EU material. Indian state-level regulatory text, in particular, is often published only on state government portals, frequently not in a format that gets well-indexed or widely reproduced elsewhere, and is amended through notifications that may never appear in the kind of aggregated legal databases that do make it into training data.

The result is a model that has seen a great deal of, say, GDPR commentary and comparatively little of a specific state's factory rules amendment from eight months ago. Asked a question about the latter, an ungrounded model does not reliably know that it does not know. It produces an answer with the same fluent confidence it would use for a question it actually has strong training coverage on — and the answer is frequently, subtly, wrong in ways that are hard to catch without independently checking the source.

The specific failure pattern we see most often

Two failure modes recur, and both are more dangerous than an obvious refusal would be, because both produce answers that read as correct.

Reasoning imported from the wrong jurisdiction. Asked about breach notification obligations, an ungrounded model will sometimes answer with GDPR-shaped reasoning — a materiality threshold, a "likely to result in a risk" qualifier — because that is the pattern most heavily represented in its training data for "data breach notification" generally. The DPDP Act has no such threshold. The answer is fluent, internally consistent, and describes a different law's obligation as if it were India's.

Confident citation of provisions that do not exist, or no longer exist in that form. Asked to cite a specific section, an ungrounded model will sometimes produce a plausible-looking citation — correct-sounding numbering, correct-sounding language — that does not correspond to an actual current provision, particularly for recently amended or recently renumbered instruments, such as the transition from the 1961 to the 2025 Income-tax Act.

Neither failure mode looks like a failure. Both look like a normal, competent answer, which is precisely what makes them dangerous to rely on without independent verification.

What grounding actually means, mechanically

A grounded system does not generate an answer from the model's general training first and check it afterward. It retrieves specific passages from a curated, current corpus — the actual Act text, the actual Rules, the actual state notification, the actual internal policy — and constrains the generated answer to what those retrieved passages actually support, with each claim in the answer traceable to a specific retrieved passage.

This changes the failure mode entirely. Where the corpus does not contain a passage supporting an answer — because the relevant state has not yet notified its rules, for instance, or because the question falls outside what has been ingested — a properly grounded system's correct behaviour is to say so, and show what it searched, rather than fall back on the underlying model's general training to fill the gap plausibly.

Why this matters more for Indian law specifically than the general AI safety conversation suggests

Most public discussion of AI hallucination treats it as a roughly uniform risk across domains. For Indian regulatory questions specifically, the risk is elevated for a structural reason: the training data asymmetry described above means an ungrounded model's confidence is least calibrated exactly where the stakes are relevant to an Indian compliance question — state-level law, recently amended provisions, and areas where Indian regulatory design deliberately differs from the more heavily-represented Western frameworks the model has seen more of.

What to actually ask when evaluating a tool

Two tests, both fast to run and both more informative than reading a vendor's marketing description of "grounded AI."

Ask it something specifically outside likely training data — a state's rules under a Code that has not yet been notified in that state, for instance. A grounded system says so. An ungrounded one, or a poorly grounded one, will often produce a plausible-sounding answer describing rules that do not yet exist.

Ask for the citation behind a specific claim, and check it. A grounded system points to a specific retrieved passage. An ungrounded one produces a citation that may or may not correspond to an actual current provision — and checking is the only reliable way to find out, because a wrong citation reads exactly like a right one until verified against the source.

The uncomfortable truth for anyone evaluating AI tools in this space: fluency is not evidence of grounding, and a demo that answers everything smoothly, with no visible refusals, is more likely to be poorly grounded than well grounded — because a genuinely grounded system, honestly built, will decline to answer some fraction of questions its corpus does not support, and a demo with zero declines is a demo that has not been asked a hard enough question yet.

AIGroundingModel risk

Written by Rohit Menon, Principal Regulatory Analyst

Part of the team that builds and maintains the Regulens obligation library and platform. If you disagree with something here, we would genuinely like to hear it — get in touch.

Everything here is how the product actually works

If the methodology in these articles matches how you think the problem should be solved, a demonstration will be a short conversation.