Label Agreement Does Not Measure Authorization

Evidence-grounded auditing for bounded LLM metadata harmonization

1Joint Georgia Institute of Technology / Georgia State University / Emory University Center for Translational Research in Data Science and Neuroimaging 2Georgia Institute of Technology

Accepted to NeurIPS 2026 (AIM Workshop)

TReNDS Center Georgia State University Georgia Tech
NeurIPS 2026
Evidence-grounded metadata harmonization and authorization pipeline
Evidence-grounded harmonization pipeline. Source evidence is extracted from read-only metadata artifacts, an LLM proposes a candidate interpretation, and a deterministic authorization gate decides whether the output is complete, entitled, and stable enough to commit.

Abstract

Neuroimaging metadata harmonization often asks whether a proposed label matches a reference label. This project shows why that is not enough: an agent can agree on a label while lacking the evidence needed to commit that label, violating the output contract, or failing when decisive evidence changes.

The system evaluates a bounded LLM workflow for harmonizing COBRE and FBIRN metadata. Read-only source readers build Evidence Cards from structure, value profiles, dictionaries, identifier overlap, lineage, and imaging row-order checks. The LLM proposes a controlled action, but a deterministic controller performs the final authorization check before anything is committed.

TL;DR

Label agreement is narrow

Matching the expected label does not prove that the agent had enough evidence, followed the contract, or made a safe commitment.

Authorization is separate

The controller checks whether required evidence exists before accepting, repairing, abstaining, or escalating a proposed output.

Downstream errors matter

Metadata mistakes that leave schema metrics untouched can still reduce neuroimaging prediction performance.

Evaluation Design

The audit keeps label agreement, action agreement, case success, unsafe commitments, coverage, and exact repair separate. It uses a real metadata panel, a synthetic paired-evidence panel, and downstream sFNC diagnosis-prediction tests.

Real panel

230 file, column, value, identity, and governance decisions from 88 base units in 16 lineage families.

Synthetic panel

90 paired cases with controlled evidence changes, including 58 pairs for controlled-error consistency.

Downstream test

Diagnosis prediction from Neuromark sFNC features for 157 COBRE rows and 311 FBIRN rows.

Key Results

The results show a recurring gap between surface-level label agreement and evidence-authorized behavior. Revealing the upstream proposal barely changes Gemma4 label agreement, while action agreement rises sharply; contract parsing alone also does not guarantee that required authorization fields are present.

0.857 -> 0.870

Gemma4 label agreement, hidden to visible proposal

0.409 -> 0.830

Gemma4 action agreement, hidden to visible proposal

0.948

Controlled-error consistency for the deterministic gate

0.000

CEC for the frozen-answer adversary after evidence changes

Contract completeness

Every contract-arm response parses as JSON, but the required authorization field is missing in 23/230 Gemma4 outputs and 77/230 Qwen3 outputs.

Downstream sensitivity

Shuffling half of the metadata-to-image row order reduces AUC by 0.13 on COBRE and 0.26 on FBIRN, even when label/schema checks remain unchanged.

System Label agreement Action agreement Case success Unsafe rate Commitment
Gemma4 hidden0.8570.4090.3870.3780.726
Gemma4 visible0.8700.8300.7650.1910.917
Gemma4 contract0.8430.7260.6430.2780.896
Qwen3 hidden0.6000.2650.2300.2000.409
Qwen3 visible0.7260.3740.2960.3520.617
Qwen3 contract0.7040.4300.3430.3480.670
Deterministic reference0.9520.9830.9520.0480.957

Reproducing the Public Panel

The public repository includes a fully synthetic panel that can be run without restricted COBRE or FBIRN data. Restricted cohort experiments require authorized NeuroMark access.

git clone https://github.com/amir-sbg/Label-Agreement-Does-Not-Measure-Authorization.git
cd Label-Agreement-Does-Not-Measure-Authorization

python -m venv .venv
source .venv/bin/activate
pip install -e ".[agent,dev]"
pytest

python scripts/run_agent.py \
  --cases data/synthetic_panel/cases.jsonl \
  --model Qwen/Qwen3-4B-Instruct-2507 \
  --interface visible \
  --output runs/demo/shards/shard_000.jsonl

python scripts/merge_evaluate.py \
  --shards runs/demo/shards \
  --oracle data/synthetic_panel/oracle.jsonl \
  --output runs/demo

BibTeX

@inproceedings{sabbaghziarani2026labelagreement,
  title     = {Label Agreement Does Not Measure Authorization},
  author    = {Sabbaghziarani, Amir and Baker, Bradley Thomas and LaGrow, Theodore J. and Plis, Sergey},
  booktitle = {NeurIPS 2026 Workshop on Agentic Intelligence for Medical Imaging and Multimodal Clinical Data},
  year      = {2026}
}