Science & validation

Research, not hype.

Speech AI should be evaluated on the people, languages, sounds and workflows where it will actually be used. This page explains what we evaluate, how, and what we will publish. For SLPs, researchers, clinical leaders and technical evaluators.

What we evaluate

Phoneme-level tasks, clinician agreement, uncertainty, recording quality.

Generic ASR benchmarks do not predict performance on the tasks Articu is used for. Our evaluation is purpose-built around the clinical workflow.

Phoneme-level classification

Substitution, omission and distortion patterns per target sound and word position.

Clinician agreement

Model output compared with independent trained-SLP judgments, with a documented adjudication methodology.

Uncertainty behavior

Abstention rate, review rate, and accuracy inside the auto-accepted and review bands.

How we evaluate

Every result carries its context.

When metrics exist, each benchmark card will include model version, language and dialect scope, sample size, age and population scope, labeler type, metric, evaluation date and a known-limitation note. Results without context do not get published.

Who labels and adjudicates

Trained SLP raters label independently; disagreements go through a documented adjudication process.

Recording-quality performance

Audio rejection rate, performance by noise level and device behavior are measured, not assumed.

Subgroup analysis

Age group, language, dialect, sound class and word position—published with each result where sample sizes allow.

Benchmark cards

Validation in progress—stated openly.

No validated metric exists yet for public release. Rather than manufacturing numbers, we publish the card structure now and fill it with real data before any autonomous feedback ships.

/r/ Initial Production ClassificationValidation in progress
Model
Articu Engine (versioned at release)
Language
en-US
Dataset
To be published with the result
Reference
Independent SLP raters + adjudication
Metric
To be published with the result
Known limitation
To be published with the result

We will publish this card with real data before enabling autonomous feedback for this target.

Clinician Agreement StudyValidation in progress
Design
Model vs. independent SLP judgment
Population
To be published with the result
Adjudication
Documented process, to be described
Metric
To be published with the result
Date
—
Known limitation
To be published with the result

We will publish this card with real data before enabling autonomous feedback for this target.

Evaluation maturity ladder

Where Articu stands today.

Each capability will state its rung on this ladder. Nothing jumps to the top without evidence.

1Literature groundedDesign decisions traceable to published researchReached
2Internal offline benchmarkEvaluation against internal test setsReached
3SLP-labeled validationComparison with trained clinician judgmentsNot yet
4Subgroup analysisPerformance by age, language, sound class, positionNot yet
5Prospective pilotEvaluation inside real clinical workflowsNot yet
6External validationIndependent replicationNot yet
7Peer-reviewed publicationPublic, citable evidenceNot yet

Versions & limitations

Model changes are versioned. Limitations are public.

Engine updates will ship with a public changelog: what changed, which benchmarks were re-run, and what limitations remain.

Claim governance

Marketing claims are matched against a claim registry: every published clinical or performance statement maps to evidence, or it is not published.

Honest status language

Internal validation, clinician-reviewed validation, preprint, peer-reviewed and independent validation are distinct labels—and we use them precisely.

Research collaboration

Validate with us.

We are interested in collaborations with clinical researchers, universities and speech programs—especially for multilingual and pediatric speech validation.

Follow our validation work.

Pilot participants see evaluation results as they develop—and help define what responsible performance reporting should look like.

Internal validation is not equivalent to independent clinical validation.