From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders
Abstract
Typed decision models answer schema-constrained questions about a text in one forward pass and return probabilities meant to be thresholded. We ask whether biomedical sentence encoders trained for retrieval are good starting points for such models. We present SBERT2S1, which converts Sentence-Transformers encoders into bi-encoder, cross-head (C) and prior-fused residual (PFR) decision models, together with BIODECIDE, a biomedical typed-decision suite, and MEDLINE-S1, 243k training decisions derived from NLM indexing. Across six parent-retriever pairs, retrieval training improves zero-shot matching of content-bearing options. After fine-tuning, its effect depends on the head: across five pairs and three training-set sizes, retrieval training significantly helps PFR, which keeps the retrieval prior, in 10 of 15 comparisons, but helps C in one and hurts it in five. A matched grid of two heads and five training objectives shows that C outperforms PFR under every objective, and that the released RLCD recipe of open System One models trails cross-entropy by 2.5-3.0 points. The deficit stems mainly from its reward normalisation, which inflates the noisy score-function term 3.6-15-fold; an unbiased leave-one-out estimator recovers most of the gap. After temperature scaling, no objective is clearly better calibrated than cross-entropy. We release the code, the MEDLINE-S1 labels and a model.
Community
This paper introduces SBERT2S1, a framework that adapts biomedical Sentence-Transformers into single-pass, schema-constrained "System One" decision models evaluated across BIODECIDE (a benchmark of grounded and clinical tasks) and trained on MEDLINE-S1 (243k decisions curated from NLM indexing). Across eleven encoders and six parent-retriever pairs, contrastive retrieval pre-training significantly improves zero-shot option matching and aids low-resource fine-tuning when paired with a novel Prior-Fused Residual (PFR) head—which achieves near-complete invariance to option ordering (0.3% flip rate vs. 15.7% for cross-encoder heads)—while standard cross-encoder heads deliver higher raw accuracy. Furthermore, theoretical and empirical analyses reveal that the popular RLCD training objective trails standard cross-entropy by 2.5–3.0 points due to an inflated reward normalization baseline, an issue resolved by a simple leave-one-out estimator, with no training objective out-calibrating temperature-scaled cross-entropy. Finally, cascading these fast ~110M base encoders with larger LLMs by escalating only the 20% least-confident decisions outperforms using either model alone, with models, code, and datasets openly released.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Benchmarking System One decision models against trained classifiers and language models for automated decision gates (2026)
- LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information (2026)
- PACT: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions (2026)
- Chinese-Jev: Bringing System One Model to Chinese-Language Tasks (2026)
- Typed Decision Models: An Early Evidence Audit and Evaluation Checklist (2026)
- Automated Regulatory Compliance Question Answering in Financial Services with Domain-Adapted Retrieval-Augmented Generation (2026)
- Locating Answer-Correctness Signals in Frozen Large Language Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.02486 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 1
pritamdeka/MEDLINE-S1
Spaces citing this paper 0
No Space linking this paper