IQAI Risk, Early Pilot

Human / LLM / IQAI Risk claim-evidence review.

A small human-in-the-loop pilot tested whether five human reviewers, multiple language models, and IQAI Risk identified similar evidence gaps across the same ten claim/evidence questions.

The useful result is not “IQAI beats humans” or “IQAI beats AI.” The pilot shows that obvious evidence gaps are recognizable across reviewers while IQAI Risk adds a repeatable review structure: claim classification, verification routing, policy boundaries, and review receipts.
01 · Study design

Same claims. Same evidence. Three review modes.

The pilot used ten claim/evidence items and compared human reviewers, general-purpose LLMs, and the IQAI Risk output using the same underlying review problem.

Human reviewers
5
MTurk participants, anonymized.
Questions
10
Claim/evidence pairs.
LLM reviewers
3
General-purpose model runs.
Support κ
0.664
Fleiss’ kappa for human support labels.
Issue κ
0.402
Fleiss’ kappa for human issue labels.
Interpretation: humans showed substantial agreement on support posture and moderate agreement on issue type. The pilot is directional evidence, not a publication-grade validation study.
02 · Human reviewer pattern

Humans agreed strongly on obvious support failures.

On most questions, the human majority converged on the same support posture: supported when the evidence was direct, weak when the evidence was partial, unsupported when the claim outran the memo, and external when outside verification was required.

Supported
2 items
Weak / Needs Review
3 items
Unsupported
3 items
Needs External Verification
2 items
The review problem itself is not exotic. Reviewers can usually recognize obvious mismatches. The challenge is turning that judgment into a repeatable control system.
03 · LLM reviewer pattern

LLMs generally recognized the same evidence gaps.

The general-purpose models largely agreed with the human majority, especially when claims clearly exceeded the provided evidence.

Agreement

Most items aligned directionally.

The models usually identified the same weak, unsupported, and external-check patterns seen by the human reviewers.

Variation

Issue labels differed more than support labels.

Models sometimes framed the same weakness differently: causal overreach, missing evidence, external verification, or overstatement.

Governance gap

Judgment alone is not a control record.

A good model answer can identify a problem, but it does not automatically create policy boundaries, review queues, or receipts.

04 · IQAI Risk output

IQAI Risk adds governance structure to the same review problem.

The pilot compares human and model judgment with a rule-governed system that separates memo evidence fit, external verification, issue type, and reviewer action.

Claim classification

Checkable fact, forecast, guarantee, comparison, recommendation, or other reliance-bearing statement.

Evidence-fit label

Supported, Weak, Unsupported, or Needs External Verification.

Verification routing

External facts can move to registry, filing, date, market, or math checks without conflating lookup with broad support.

Review receipt

What was checked, flagged, accepted, modified, escalated, and left unresolved.

05 · Calibration boundary

The most important finding is a policy question.

The pilot exposed a meaningful boundary between Weak / Needs Review and Unsupported.

Forward-looking guarantee with no direct evidence

“The new process ensures incidents will be escalated quickly.”

Humans and LLMs often moved this toward Unsupported because the memo proved only that a process was updated, not that future incidents would be escalated quickly. IQAI Risk can be tuned so this pattern is treated as Unsupported rather than merely Weak.

Partial evidence with overbroad wording

“The engagement improved alignment across all departments.”

If the memo contains meeting records or partial survey results but does not support the broad “all departments” wording, Weak / Needs Review may be the appropriate posture rather than Unsupported.

This is where a rule-governed product becomes useful: the organization can decide its policy boundary and apply it consistently instead of leaving every reviewer to improvise.
06 · What the pilot supports

Repeatable review infrastructure is the product.

The pilot does not support a claim that IQAI is universally more accurate than humans or LLMs. It supports a narrower and more useful claim: the same review judgments can be turned into repeatable infrastructure.

Consistent labels

Policy can be encoded.

The organization can define how evidence gaps map to review states.

Verification separation

Outside facts stay outside.

Registry, filing, market, date, and math checks do not automatically certify broader claims.

Human authority

Review remains explicit.

Humans can accept, revise, escalate, or reject the system’s posture.

Receipt

The review becomes reconstructable.

The result can be preserved as a claim/evidence record rather than a transient judgment.

Pilot conclusion

Humans and LLMs can see the gap. IQAI Risk turns the gap into a governed record.

That is the practical value: evidence fit, verification routing, policy calibration, human review, and receipts before reliance.