LLMs are brilliant at sounding right. That is also the problem.
In production, a model can return a clean JSON object, a confident summary, or a polished support reply and still be wrong in ways that are expensive, embarrassing, or quietly destructive. A refund policy gets invented. A tool name gets hallucinated. A claim slips through without evidence. The output parses fine, but the business logic behind it starts to drift.
Here’s an uncomfortable question. Of every LLM call your application makes today, how many actually need an LLM?
I was running multiple AI agents inside one application. Some were doing real work, the kind that justifies a few hundred billion parameters. Others were doing things like: is this JSON shaped correctly, does this tool name actually exist, does this sentence match the document I gave you. That’s not reasoning. That’s a lookup. And I was paying model prices for a lookup, every single time, at scale, without noticing.
So let’s talk about the actual problem, LLM hallucination in production systems and then the boring, unglamorous, open-source fix I built for it.
What is LLM hallucination, actually?
Everyone talks about hallucination like it’s a bug that occasionally shows up. It’s not a bug. It’s a feature of how these models work, they’re built to sound right, not to be right, and most of the time those two things happen to line up. Most of the time.
Here’s what it looks like when they don’t, in real production systems:
- a support bot invents a refund policy that was never written, and says it with total confidence,
- an agent calls a tool named something like rm -rf because it guessed the function name and nobody stopped it,
- a summarizer cites your own source document for a claim the document never made,
- structured JSON output that parses perfectly clean and is wrong in every field that matters.
Notice the pattern? None of these come with a warning label. The model doesn’t pause, doesn’t hedge, doesn’t go “I’m 60% sure about this.” It just says it, smooth, fluent, and occasionally fictional. That’s the whole risk in one sentence: the failure mode looks identical to the success mode. You can’t eyeball your way out of this one. You need something scoring the output before a human or another system ever sees it.
So I built the thing I actually needed, not the thing that sounds impressive
I built hallx, a lightweight, open-source hallucination-risk scoring library for production LLM pipelines. It sits between your model’s output and whatever system is about to act on that output: a database, a customer, another agent, whatever’s downstream. It doesn’t make the model smarter. It just doesn’t let a bad answer walk through unchecked, which, it turns out, is most of the actual problem.
pip install hallx
from hallx import Hallx
result = Hallx(profile="balanced").check(
prompt="Refund policy",
response="Refunds are allowed within 30 days.",
context=["Refunds are allowed within 30 days of purchase."],
)
print(result.confidence, result.risk_level) # 0.93 low
print(result.issues, result.recommendation) # [] {'action': 'proceed', ...}
Five checks run before anything is trusted: is the structure valid, is the answer consistent across repeated generations, is it grounded in the context you actually gave it, which specific claims are invented, and for agents are the tool calls even real. None of this is exciting. That's kind of the point. The exciting stuff is usually where the bugs live.
The part that actually saves you money, pay attention here
This is the bit most “AI safety” content skips entirely, because it’s not sexy enough to tweet about: most of what you’re checking doesn’t need a model at all.
Checking a JSON schema doesn’t need an LLM. It needs jsonschema, and it runs in microseconds. Checking if a sentence roughly matches your context doesn't need an embedding call every time either, rapidfuzz string similarity gets you most of the way there, for free, with zero GPU and zero network hop. That's why hallx's core install has exactly three dependencies jsonschema, rapidfuzz, requests and, I want to be very clear about this, none of them are "call another model to check the first model." The heavy stuff, real NLI cross-encoders, LLM-as-judge scoring, is there. It's just opt-in, behind the hallx[nli] extra. You reach for it when the cheap check genuinely can't tell you enough, not by default.

Do the math on that at enterprise scale. Thousands of agent steps a day, most of them repetitive, most of them checkable with a boolean instead of an API call. That’s not a rounding error on your token bill. That’s the difference between a bill your CFO tolerates and one that gets you a meeting you don’t want to be in.
Wait, what is NLI, and why does hallx even care?
Quick detour, because I used the term LocalNLIChecker a few paragraphs back like everyone already knows what that means, and that's a bad habit to have in a blog that's supposed to be readable.
NLI stands for Natural Language Inference. It’s a specific, well-studied task in NLP where a model looks at two pieces of text, a premise (your context) and a hypothesis (the claim you’re checking) and decides whether the premise entails, contradicts, or is neutral toward the hypothesis. Not “sounds similar.” Not “shares a lot of the same words.” Actually entails, as in, if the premise is true, does the hypothesis logically follow.
That distinction matters more than it sounds like it should. Fuzzy string matching, the kind hallx uses by default, is really good at catching things like “these two sentences are basically the same wording.” It’s much worse at catching things like: “The refund window is 30 days” versus “The refund window is not 30 days.” Those two sentences are extremely similar by any string-distance metric, swap one word and you’ve flipped the entire meaning but a fuzzy matcher has no concept of negation. An NLI model does. It reads “not” and understands that changes everything, because that’s the actual task it was trained on.
That’s what LocalNLIChecker gives you, a real cross-encoder model (cross-encoder/nli-deberta-v3-base, if you're curious) doing genuine entailment/contradiction judgment instead of approximating it with string overlap. It's heavier than rapidfuzz, which is exactly why it's opt-in behind the hallx[nli] extra rather than sitting in the core install. You don't want to pay that cost on every single check, you want it available for the claims where fuzzy matching genuinely isn't enough, which, to be fair, is a smaller slice of your traffic than you'd expect.
Claim-level grounding: telling you exactly which sentence is lying
A confidence score of 0.7 is technically information. It’s just not useful information, it doesn’t tell you where the problem is, only that there is one, like a smoke alarm that won’t tell you which room’s on fire. So hallx goes sentence by sentence, extracting each claim, mapping it to the closest evidence in your context, and labeling it:
from hallx.attribution import check_claim_grounding
response = "The Eiffel Tower is in Paris. The Eiffel Tower is also in Berlin."
context = ["The Eiffel Tower is located in Paris, France."]
result = check_claim_grounding(response, context)
for claim in result.claims:
if claim.status != "filtered":
print(f"[{claim.status}] {claim.text}")
[supported] The Eiffel Tower is in Paris.
[unsupported] The Eiffel Tower is also in Berlin.
Somewhere, apparently, there’s a version of Paris in Berlin. Good to know. Good to catch, too, before that sentence ends up in a customer-facing summary and someone in support has to explain German geography to a confused user.

The result object gives you score, supported_count, weak_count, unsupported_count, and per-claim similarity, evidence_index, and evidence_snippet, so you're not just getting a verdict, you're getting a paper trail. Want real entailment instead of fuzzy string matching? Swap in a local NLI cross-encoder, or route it through an LLM-as-judge using any of the eight provider adapters hallx ships with:
from hallx import OpenAIAdapter
from hallx.judge import GroundingJudge
judge = GroundingJudge(llm_adapter=OpenAIAdapter("gpt-4o-mini", api_key="..."))
result = check_claim_grounding(response, context, verifier=judge)
Same logic covers tool calls in agentic loops, every invocation gets checked against its schema before it executes, so a hallucinated function name gets caught in review, not in your incident channel:
from hallx import Hallx
tools = {
"get_weather": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
"additionalProperties": False,
},
}
checker = Hallx()
agent_output = [
{"function": {"name": "get_weather", "arguments": '{"city": "Paris"}'}},
{"name": "rm -rf", "arguments": "{}"}, # hallucinated tool name
]
result = checker.check_tool_call(agent_output, tools)
print(result.score) # 0.25
print(result.verdicts) # ok / unknown_tool
if result.recommendation["action"] == "block":
print("Regenerate before invoking any tool.")
What’s actually shipped, not the roadmap, the real thing
Anyone can write a README full of plans. Here’s what’s live, tested, and running in the current release:
- Three-axis core scoring schema validation (JSON Schema Draft 7, plus a custom null-injection detector), consistency across repeated generations, grounding against your context.
- Claim-level grounding sentence-level attribution with span offsets, filler filtered out, every real claim labeled supported, weak, or unsupported.
- Pluggable verifiers LocalNLIChecker for real cross-encoder entailment, or GroundingJudge, an LLM-as-judge hardened against prompt injection in its own scoring prompt, because apparently even your safety layer needs a safety layer.
- Tool-call validation every agent tool call checked against its schema, verdicts of ok, unknown_tool, invalid_arguments, malformed, invalid_definition, with a proceed/fix/block recommendation attached.
- Profiles: fast, balanced, strict tune consistency runs and penalties, or bring your own weights if you don't trust my defaults either.
- Strict mode and SLO gating: Hallx(strict=True) raises HallxHighRiskError on high-risk output instead of quietly letting it through, and assert_safe() lets you set hard thresholds of your own.
- Feedback and calibration: record_outcome() logs human-reviewed verdicts to a local SQLite store, calibration_report() turns that history into an actual recommended confidence threshold via F1 search, instead of a number I made up.
- Sync and async, properly: check / check_async, check_claim_grounding / check_claim_grounding_async, parallel embedding calls via asyncio.gather, built for services under real concurrency, not a demo notebook.
- Eight provider adapters: OpenAI, Anthropic, Gemini, OpenRouter, Perplexity, Grok, HuggingFace, Ollama, or bring your own plain Python callable if you’re running something nobody’s heard of yet.
Where this actually gets used
If you’re wondering whether hallucination detection applies to your stack, here’s where hallx tends to show up in practice:
- RAG and grounded question-answering: verifying every sentence of a generated answer against retrieved documents, and regenerating when claim-grounding flags something unsupported.
- Agentic tool-calling loops: validating tool names and arguments before a handler ever executes them, so a hallucinated function call gets blocked, not run.
- Machine-consumed / structured output: schema and null-injection checks so a downstream database or API only ever receives valid, sane input.
- Citation integrity in generated content: catching fabricated sources, fake DOIs, and unverifiable URLs before a summary goes out under your name.
- Support and helpdesk assistants: verifying policy answers against the actual policy context before they reach a customer.
- Research and news synthesis: where ungrounded claims across repeated generations are exactly what damages trust in the product.
- CI/CD for LLM behavior: running hallx as an automated evaluator in an eval harness, with the feedback loop tightening risk thresholds over time.
- High-concurrency async services: check_async and parallel embeddings so guardrails don't become the bottleneck.
Where that leaves you
I did not want hallx to be another “AI safety” library that looks good in a diagram and disappears in production. The goal was simpler: make hallucination risk measurable, make the result usable, and keep the implementation light enough that teams will actually adopt it.
That is the product philosophy behind hallx: score the response, surface the issue, and give the system a decision it can act on. No drama. No inflated ceremony. Just a guardrail that does its job.
If you are building production LLM systems and tired of trusting outputs on vibes alone, hallx is meant to help.
And yes, it is also one of the quieter ways to reduce token spend, because the cheapest token is the one you never had to waste.
pip install hallx
pip install "hallx[nli]" # if you want local NLI too
Repo’s at dhanushk-offl/hallx, docs at hallx.pages.dev. Go find where the scoring is wrong. I'd rather you tell me than find out from a production incident.