2026-10-04
Evidence supplement
Measuring the Measure: Evidence and Limits
Text-free outcome counts, recomputation code and provenance limits for our Malay LLM safety case study, stating what the published evidence can and cannot establish.
Oaica is a safety-focused AI research company. We publish the methods, the evidence and the limits behind evaluating AI systems — across languages, models and deployments.
How safety is measured, what a benchmark score supports, and where inference has to stop.
Evaluation methods that hold up outside English, starting from Malay and extending to other regional languages.
Reasoning modes, refusal behaviour and robustness under matched, reproducible conditions.
Running capable models on your own hardware, with the limits of each configuration stated plainly.
Text-free outcome counts, recomputation code and provenance limits for our Malay LLM safety case study, stating what the published evidence can and cannot establish.
One changed prediction can move a Malay LLM safety subtask by up to 9 points. A study of the SEA-HELM Malay safety track, an Indonesian cross-language probe, and a checklist for any language.
The Oaica 35B-A3B Malay model, its safety-tuned variant and two experimental specialists: what we measured, the trade we report, and what the copy may say.
Oaica's 35B-A3B Malay model runs on your own hardware, from small GPUs to datacenter servers, with our .oqm format and a statement of what we do not claim.
How far a Malay AI safety score can be trusted. One changed prediction can move a Malay LLM safety subtask by up to 9 points. What the SEA-HELM Malay safety benchmark can and cannot resolve, with an Indonesian cross-language probe, by Oaica.
Oaica, a safety-focused AI company: one changed answer can move a Malay AI safety sub-score by up to 9 points. On-premises Malay AI and safety audits.
Oaica researches AI safety and evaluation. We develop and study language models, and publish the foundations behind them: how a benchmark score is produced, what its intervals support, and what remains unverified. Our first programme is Malay-language safety; the methods are built to extend to other languages and models.
De-identified evaluation tables and item-leverage worksheets are available on request. Nothing on this site certifies any model as safe.