๐ Highlights
ASCRIBE is a structured clinical reasoning framework for Thai SOAP-note generation that extracts atomic transcript facts, classifies their clinical significance, and conditions the final SOAP note on that trace.
The strongest result is not a single metric but a pattern: ASCRIBE improves frontier prompting over standard and CoT prompting, while ASCRIBE-style GRPO rewards improve clinically aligned metrics for a smaller Gemma-4-E4B-it model.
The main caveat is that the evaluation is still narrow: small outpatient datasets, Thai only, one SLM, LLM judges validated on limited encounters, and no real physician user study of generated notes.
- Core claim: structured atomic and significance-based reasoning is more reliable for Thai clinical SOAP-note generation than direct prompting or free-form CoT.
- Best frontier prompting row: GPT-5.4 with ASCRIBE reaches 95.2/59.6/71.2 SP/CE/CE-Impact on ICOPD and 98.3/62.6/74.6 on ThaiClinicBench.
- Best synthetic-only GRPO row: Gemma-4-E4B-it with ASCRIBE reward and S-ICOPD reaches 86.0/64.9/75.7 on ICOPD and 89.6/69.7/77.9 on ThaiClinicBench.
- Important limitation: ASR transcripts reduce CE and SP for all evaluated models, with GRPO-trained models generally showing larger drops than frontier LLMs.
๐ฏ Introduction
The paper targets automatic SOAP-note generation for Thai physician-patient encounters. The clinical motivation is documentation burden: comprehensive notes are important for patient safety but time-consuming, especially in regions with low doctor-to-patient ratios. The operational pipeline is transcribe-then-summarize: ASR converts the clinical encounter into text, then a model converts the transcript into a Subjective, Objective, Assessment, Plan note.
The authors argue that existing approaches leave two gaps. First, clinical summarization models trained primarily on English datasets remain exposed to omission, hallucination, and out-of-domain weakness. Second, Thai is low-resource for this task: the main body states that no public Thai dataset supporting SOAP note generation existed to their knowledge. Thai conversational nuance, dialectal clinical terms, and ASR errors further complicate deployment.
ASCRIBE addresses the reasoning gap by replacing free-form CoT with a physician-inspired trace over atomic clinical facts and significance labels. It addresses the data gap with S-ICOPD, a synthetic Thai training corpus seeded from real clinical notes, and ThaiClinicBench, a de-identified benchmark of 44 real Thai clinic encounters with physician-written SOAP notes.
๐ฌ Methodology
The method starts from a transcript x and asks the model to produce an intermediate trace A before the final SOAP note. Each trace item is an atomic fact fi, a short supporting evidence span ei from the transcript, and a predicted clinical-significance level si. This is meant to force the model to identify grounded, clinically relevant units rather than summarizing directly.
Clinical significance is not treated as generic salience. The paper defines six ordered levels co-designed with certified physicians: Diagnosis-impacting, Treatment-impacting, Required, Follow-up-relevant, Future-reference, and Irrelevant. The first three map closely to facts that support assessment, management, and required documentation; the lower levels preserve longitudinal but less immediately actionable details or filter out administrative and redundant content.
For smaller-model training, the paper compares LoRA SFT and GRPO. In SFT, the model is trained either to generate standard SOAP notes or to generate structured reasoning prepended to the target. In GRPO, the model receives five scaled rewards: atomic fact reward, clinical significance reward, format reward, completeness reward using ClaimEntailment, and factual grounding reward using SOAPPrecision.
The synthetic data pipeline adapts an intent-graph approach. It seeds generation from SOAP notes, plans an intent path through a weighted directed graph learned from real conversations, and uses doctor and patient agents to generate dialogue section by section. LLM judges validate whether intents are fulfilled and whether sections satisfy adapted Thai clinical conversation rules.
Table 1
Caption: Clinical significance levels used for categorizing atomic facts. Levels are listed from most to least significant.
| Level | Description |
|---|---|
| Diagnosis-impacting | Essential information supporting the diagnosis or assessment of the current visit (e.g., chief complaint, symptoms, clinical findings, disease progression, physical examination, signs, laboratory or imaging findings). |
| Treatment-impacting | Details regarding or influencing the management of the current visit (e.g., medications, procedures, follow-up plans, treatment recommendations). |
| Required | Patient information required by clinical standards, regardless of the current complaint (e.g., comorbidities, drug allergies, smoking/alcohol history). |
| Follow-up-relevant | Details not essential to the current visit but useful for subsequent follow-ups (e.g., advices given, referrals, treatment goals). |
| Future-reference | Details useful for future reference but unlikely to affect the current episode of care (e.g., irrelevant family history, social history, unrelated medical history). |
| Irrelevant | Does not meaningfully contribute to current or future clinical decision-making (e.g., redundant statements, vague complaints, hyper-specific context, administrative details, physician explanations). |
Why It Matters: The significance taxonomy is the paper's core clinical structure and is used in both prompting and reward design.
equation 1
Meaning: Defines the ASCRIBE trace as extracted atomic facts, transcript evidence spans, and clinical-significance labels.
equation 2
Meaning: Makes the final SOAP note conditional on both transcript and structured reasoning trace.
display equation 5
Meaning: Computes atomic precision for generated facts.
display equation 6
Meaning: Computes atomic recall against reference facts.
๐ Experiments
The main datasets are ICOPD, ThaiClinicBench, and S-ICOPD. ICOPD is a private corpus of 400 consented Thai outpatient encounters split into 300 training, 50 validation, and 50 test examples. ThaiClinicBench contains 44 manually de-identified and consented real OPD encounters used as an out-of-distribution test set. S-ICOPD contains synthetic transcripts generated from ICOPD notes, with 295 training and 50 validation synthetic transcripts after excluding 5 unusable training cases.
The evaluation uses ROUGE-L and BLEURT for lexical and semantic similarity, plus physician-aligned LLM-judge metrics. ClaimEntailment measures completeness as recall over reference atomic facts; CE-Impact restricts CE to Diagnosis-impacting, Treatment-impacting, and Required facts; SOAPPrecision measures factual grounding by checking whether generated SOAP statements are supported by the transcript. Higher values indicate better performance for all metrics, and SP, CE, and CE-Impact are on a 0-100 scale.
For zero-shot prompting, ASCRIBE improves Gemini 3.1 Pro over standard prompting by +3.4/+2.2/+3.7 SP/CE/CE-Impact on ICOPD and +2.6/+6.8/+9.1 on ThaiClinicBench. GPT-5.4 shows larger gains of +4.7/+5.1/+7.2 on ICOPD and +5.8/+6.9/+10.3 on ThaiClinicBench. CoT gives smaller CE gains and sometimes reduces SP or conventional metrics.
For Gemma-4-E4B-it, SFT tends to preserve stronger ROUGE-L, but GRPO with clinically grounded rewards improves SP, CE, and CE-Impact. Among real-data GRPO variants, ASCRIBE reward with ASCRIBE prompt reaches 87.1/62.1/71.2 on ICOPD and 89.4/64.0/72.1 on ThaiClinicBench. Synthetic-only ASCRIBE GRPO improves CE and CE-Impact further to 64.9/75.7 on ICOPD and 69.7/77.9 on ThaiClinicBench, while mixed Real+Synth gives the highest ThaiClinicBench CE of 70.0 but lower SP.
- Central benchmarks: ICOPD test split, ThaiClinicBench, and ASR-transcript ThaiClinicBench robustness evaluation.
- Central baselines: standard prompting, CoT prompting, ASCRIBE prompting, SFT, ROUGE-L GRPO, BGE-M3 GRPO, ClinicEval GRPO, and ASCRIBE GRPO.
- Metric direction: all metrics are higher-is-better; SP, CE, and CE-Impact are reported on a 0-100 scale.
tab:main_results
Caption: Comparison of ASCRIBE prompting and fine-tuning against baseline methods on the ICOPD test split and ThaiClinicBench. Ours denotes the corresponding ASCRIBE prompt or reward, Real denotes training on ICOPD, and Synth denotes training on S-ICOPD. Higher values indicate better performance for all metrics. Underline marks the best result within each model/training block, and bold marks the best result within each approach (frontier LLM prompting vs.\ Gemma-4-E4B-it fine-tuning). SP, CE, and CE-Impact are reported on a 0--100 scale.
| Model / Training | GRPO Reward | Training Data | Prompt | ICOPD SP | ICOPD CE | ICOPD CE-Impact | ICOPD ROUGE-L | ICOPD BLEURT | ThaiClinicBench SP | ThaiClinicBench CE | ThaiClinicBench CE-Impact | ThaiClinicBench ROUGE-L | ThaiClinicBench BLEURT |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemini 3.1 Pro | -- | -- | Standard | 83.5 | 52.6 | 63.3 | 0.3467 | 0.3366 | 86.6 | 51.6 | 58.8 | 0.3390 | 0.2488 |
| Gemini 3.1 Pro | -- | -- | CoT | 81.0 | 53.2 | 64.3 | 0.3387 | 0.3177 | 84.7 | 52.0 | 59.8 | 0.3205 | 0.2430 |
| Gemini 3.1 Pro | -- | -- | Ours | 86.9 | 54.8 | 67.0 | 0.3406 | 0.3220 | 89.2 | 58.4 | 67.9 | 0.3170 | 0.2525 |
| GPT-5.4 | -- | -- | Standard | 90.5 | 54.5 | 64.0 | 0.3413 | 0.3299 | 92.5 | 55.7 | 64.3 | 0.3023 | 0.2623 |
| GPT-5.4 | -- | -- | CoT | 90.9 | 58.6 | 68.2 | 0.3169 | 0.3194 | 93.3 | 58.6 | 68.8 | 0.2809 | 0.2545 |
| GPT-5.4 | -- | -- | Ours | 95.2 | 59.6 | 71.2 | 0.3032 | 0.3237 | 98.3 | 62.6 | 74.6 | 0.2781 | 0.2627 |
| GRPO | ClinicEval | Real | Standard | 83.5 | 61.4 | 70.3 | 0.2293 | 0.3143 | 86.9 | 63.7 | 72.3 | 0.2052 | 0.2396 |
| GRPO | ClinicEval | Real | Ours | 84.8 | 61.8 | 71.0 | 0.2551 | 0.2979 | 88.3 | 63.5 | 71.8 | 0.2323 | 0.2467 |
| GRPO | Ours | Real | Ours | 87.1 | 62.1 | 71.2 | 0.2773 | 0.2898 | 89.4 | 64.0 | 72.1 | 0.2487 | 0.2450 |
| GRPO | Ours | Synth | Ours | 86.0 | 64.9 | 75.7 | 0.2478 | 0.2786 | 89.6 | 69.7 | 77.9 | 0.2143 | 0.2280 |
| GRPO | Ours | Real+Synth | Ours | 83.2 | 64.4 | 75.1 | 0.2306 | 0.3203 | 85.2 | 70.0 | 77.0 | 0.2105 | 0.2509 |
Why It Matters: This selected table preserves the core prompting, reward-ablation, and synthetic-training comparisons most relevant to builders.
fig:atomic-medsig-diagram

Caption: ASCRIBE's SOAP generation pipeline. Atomic facts are extracted from the transcript $x$ and classified by their clinical significance to form the reasoning trace $A$, which conditions generation of the final SOAP note $y$.
Why It Matters: This is the clearest architecture view: it shows ASCRIBE turning a Thai clinical transcript into grounded atomic facts, significance labels, and then a structured SOAP note.
fig:rlvr

Caption: GRPO training pipeline. The generated reasoning trace $A$ and SOAP note from the policy LLM are scored by the five rewards of Section~sec:rlvr, which jointly drive the GRPO update. Certain rewards require reference facts whose generation details are in Section~sec:reference_summary_gen.
Why It Matters: It connects the method to training: the policy model is rewarded not only for the final note, but also for fact extraction, significance attribution, format, completeness, and factual grounding.
fig:asr

Caption: Human-annotated vs.\ ASR transcript performance on ThaiClinicBench for CE and SP metrics. R+S refers to mixing real and synthetic training data and CEval refers to the ClinicEval GRPO reward. Blue indicates settings that use the ASCRIBE prompt and, where applicable, the ASCRIBE GRPO reward. Labels denote the score drop from human-annotated to ASR transcripts. Our structured reasoning and synthetic training achieve the highest scores on ASR in several settings despite larger drops from human transcripts.
Why It Matters: This is the deployment stress test: ASR transcripts reduce both completeness and factual grounding, and the figure shows where ASCRIBE variants remain stronger under noisy input.
๐ฎ Conclusion
The paper concludes that ASCRIBE improves Thai SOAP-note generation by explicitly modeling atomic clinical facts and clinical significance before generating the note. As a prompt, it consistently improves frontier LLMs over standard and CoT prompting on clinical LLM-judge metrics. As a training paradigm, it gives the strongest GRPO reward formulation for factual grounding and completeness among the tested Gemma-4-E4B-it variants.
The practical implication is that small models can become more clinically useful when optimized around verifiable intermediate reasoning and clinical-quality rewards. The synthetic-data results are also important: ASCRIBE training with synthetic encounters can be competitive without adding more real clinical conversations, which is useful in settings constrained by data governance and patient privacy.
๐ ๏ธ Future Research Improvements
The most direct next research step is robustness to noisy ASR transcripts. The paper shows that ASR errors degrade CE and SP for all models and that GRPO-trained models generally suffer larger drops than frontier LLMs, so future training should explicitly incorporate ASR noise or transcript uncertainty.
A second direction is better real/synthetic data mixing. The main body reports that synthetic-only training can be strong, but naive Real+Synth mixing gives inconsistent improvements: it lowers CE in every SFT setting and lowers SP in mixed GRPO settings despite high CE.
The evaluation should also be broadened. The paper itself notes the need for more encounters, additional languages or clinical settings, more SLMs, LLM-judge validation against more physicians, and real user studies with clinicians reviewing generated notes.
- Train and evaluate with ASR-corrupted transcripts rather than treating ASR as only a post-hoc robustness test.
- Study curriculum, weighting, or domain-adaptation methods for mixing real and synthetic clinical conversations.
- Validate ClaimEntailment and SOAPPrecision with multiple clinicians and report inter-rater agreement.
- Extend the benchmark beyond Thai outpatient encounters before claiming broader clinical generality.
๐ญ Potential Industry Use Scenarios
The most credible product scenario is an on-premise Thai medical scribe for outpatient clinics. The paper explicitly motivates SLMs because they can be deployed on-premise, and ASCRIBE is designed for clinical environments where data governance constrains access to external LLMs and real patient data.
A second use case is internal QA for clinical note generation systems. The same atomic-fact and significance-label structure could help audit whether generated notes omit diagnosis-impacting, treatment-impacting, or required facts, and whether generated SOAP statements are supported by the transcript.
A third scenario is synthetic data generation for low-resource clinical NLP. S-ICOPD suggests that real-note-seeded synthetic transcripts can support SFT and GRPO experiments when real transcripts are scarce, though the paper's caveats make clinical validation and leakage checks mandatory.
- Thai outpatient SOAP-note drafting from human or ASR transcripts.
- Clinical summarization QA dashboards centered on omitted facts and unsupported SOAP statements.
- Synthetic clinical-dialogue generation for low-resource medical NLP experimentation.
- On-premise small-model deployment where sending clinical transcripts to external APIs is constrained.
๐ฌ Critical Analysis
The paper's main strength is alignment between method and measurement. ASCRIBE does not merely ask for better notes; it defines clinically meaningful intermediate units, uses them in prompting, trains with rewards tied to those units, and evaluates with metrics meant to capture completeness and factual grounding. That makes the system design coherent for builders.
The strongest empirical signal is that lexical similarity and clinical quality diverge. SFT and ROUGE-L-optimized GRPO can maintain high ROUGE-L without corresponding CE gains, while ClinicEval and ASCRIBE GRPO improve clinical metrics at lower ROUGE-L. This supports the paper's argument that conventional summarization metrics are not enough for clinical note quality.
The main reproducibility and validity questions are clinical rather than purely technical. The LLM judges are central to both training and evaluation, and although the main body says they were validated for physician alignment, the limitations state that validation used limited encounters and a single board-certified physician. Builders should treat the reported metrics as promising but not sufficient for safety-critical deployment.
The synthetic data story is useful but not settled. Synthetic-only GRPO is impressive on CE and CE-Impact, but mixed real/synthetic training is inconsistent, and the authors hypothesize distributional differences between real conversations and synthetic transcripts. That makes S-ICOPD a strong starting point for research, not a drop-in replacement for deployment evaluation on real clinical workflows.