Human / LLM / IQAI Risk claim-evidence review.
A small human-in-the-loop pilot tested whether five human reviewers, multiple language models, and IQAI Risk identified similar evidence gaps across the same ten claim/evidence questions.
Same claims. Same evidence. Three review modes.
The pilot used ten claim/evidence items and compared human reviewers, general-purpose LLMs, and the IQAI Risk output using the same underlying review problem.
Humans agreed strongly on obvious support failures.
On most questions, the human majority converged on the same support posture: supported when the evidence was direct, weak when the evidence was partial, unsupported when the claim outran the memo, and external when outside verification was required.
LLMs generally recognized the same evidence gaps.
The general-purpose models largely agreed with the human majority, especially when claims clearly exceeded the provided evidence.
Most items aligned directionally.
The models usually identified the same weak, unsupported, and external-check patterns seen by the human reviewers.
Issue labels differed more than support labels.
Models sometimes framed the same weakness differently: causal overreach, missing evidence, external verification, or overstatement.
Judgment alone is not a control record.
A good model answer can identify a problem, but it does not automatically create policy boundaries, review queues, or receipts.
IQAI Risk adds governance structure to the same review problem.
The pilot compares human and model judgment with a rule-governed system that separates memo evidence fit, external verification, issue type, and reviewer action.
Claim classification
Checkable fact, forecast, guarantee, comparison, recommendation, or other reliance-bearing statement.
Evidence-fit label
Supported, Weak, Unsupported, or Needs External Verification.
Verification routing
External facts can move to registry, filing, date, market, or math checks without conflating lookup with broad support.
Review receipt
What was checked, flagged, accepted, modified, escalated, and left unresolved.
The most important finding is a policy question.
The pilot exposed a meaningful boundary between Weak / Needs Review and Unsupported.
Forward-looking guarantee with no direct evidence
Humans and LLMs often moved this toward Unsupported because the memo proved only that a process was updated, not that future incidents would be escalated quickly. IQAI Risk can be tuned so this pattern is treated as Unsupported rather than merely Weak.
Partial evidence with overbroad wording
If the memo contains meeting records or partial survey results but does not support the broad “all departments” wording, Weak / Needs Review may be the appropriate posture rather than Unsupported.
Repeatable review infrastructure is the product.
The pilot does not support a claim that IQAI is universally more accurate than humans or LLMs. It supports a narrower and more useful claim: the same review judgments can be turned into repeatable infrastructure.
Policy can be encoded.
The organization can define how evidence gaps map to review states.
Outside facts stay outside.
Registry, filing, market, date, and math checks do not automatically certify broader claims.
Review remains explicit.
Humans can accept, revise, escalate, or reject the system’s posture.
The review becomes reconstructable.
The result can be preserved as a claim/evidence record rather than a transient judgment.
Humans and LLMs can see the gap. IQAI Risk turns the gap into a governed record.
That is the practical value: evidence fit, verification routing, policy calibration, human review, and receipts before reliance.