@misc{wong2026-malay-llm-safety-score,
author = {Wong, Sam},
title = {Malay AI Safety Evaluation: A Case Study with Oaica’s Fine-Tuned 35B-A3B Model, and an Indonesian Cross-Language Probe},
year = {2026},
howpublished = {Oaica Research},
url = {https://research.oaica.com/2026/10/malay-llm-safety-score/}
}What the SEA-HELM Malay safety benchmark measures, an Indonesian cross-language probe, and what they cannot resolve
Oaica (oaica.com)
1 October 2026
On the Malay safety track of SEA-HELM (Southeast Asian Holistic Evaluation of Language Models), a single changed prediction can move a subtask score by up to 9 points and the composite by 1.5. In this case study, selected evaluation settings produced score differences comparable in size to variation among closely related historical runs. We ask what this track, a public regional benchmark, can and cannot resolve. Using two checkpoints of Oaica 35B-A3B Malay, a Malay large language model (LLM) fine-tuned from the open-weight base Qwen3.6-35B-A3B, and 20 training variants retained from one lineage’s final runs and several earlier lineages, we document how observed scores vary with evaluation choices and examine examples where measurement effects complicate interpretation.
The choice of aggregation moved one checkpoint by 5.3 points, a change in request concurrency shifted the three cultural safeguard subtasks of one model by an amount worth about 1.8 composite points, and switching reasoning on lowered the composite by 1.5 and 2.3 points, both with 95% intervals on the paired difference that include zero. By comparison, 12 training variants of one lineage spanned 3.3 points, any two of them differed by 1.2 points on average, and a single run’s 95% interval was about ±6 points. We then assess the benchmark itself: toxicity performance remained low in the evaluated runs, small class sizes give individual items high leverage, and overlapping published intervals do not establish which model differences are statistically distinguishable. These observations do not identify defects in the labels or establish a universal performance ceiling. A cross-language probe in Indonesian, a language the model was not tuned for, reproduces the same measurement behaviour (Section 5.4). Training methods are out of scope. A text-free evidence document records the available evidence and unresolved provenance.
Keywords: Malay LLM safety benchmark; LLM safety evaluation; Malay (Bahasa Melayu); Indonesian (Bahasa Indonesia); cross-language transfer; multilingual and culturally grounded safety; Southeast Asian languages; SEA-HELM; SEA-SafeguardBench; confidence intervals; annotation disagreement; benchmark reliability; leaderboard ranking uncertainty; reasoning models; open-weight models.
It does not rank models, and it does not claim that any model is safe, is the safest, or is safer than another. It does not describe training methods or data sources, and identifies the checkpoints by name, base model, size, architecture and serving characteristics only (Section 6). Its evidence comes from one primary language, Malay, with a single cross-language probe in Indonesian (Section 5.4), one benchmark family and, for the safety findings, one base model (Section 8).
This paper assesses what a single published number can support, not the intent or quality of the work behind it. SEA-HELM and SEA-SafeguardBench are among very few public instruments for culturally grounded safety in the region, and the leaderboard maintainers’ practice of publishing intervals over repeated runs [2] is what made the interval analysis in Section 5 possible. We offer our evaluation tables and item-leverage worksheets on request.
Evidence notice. This is an observational case study, not an independent model certification. Historical scores are distinguished from recomputed runs. The evidence document supplies text-free counts, exact score-recomputation code and an explicit list of missing provenance. Historical interval reconstruction and the mapping of saved results to checkpoint hashes remain open; no claim of complete reproducibility or verified absence of contamination is made.
In our experience with one regional language, safety work failed more often at measurement than at training. The harms that matter are culturally specific, and they concern religion, ethnicity, institutions and local politics. The public benchmarks that measure them in Malay are small [2], and their labels necessarily encode one annotation convention, which other competent annotators may not share [7]. SEA-HELM and SEA-SafeguardBench are valuable public instruments [1, 7]; the scorer comparison in Section 2 and the interval analysis in Section 5 are possible because the leaderboard’s maintainers document the scorer and publish intervals over repeated runs [2], and the assessment in Section 5 draws on the annotator agreement that the source benchmark reports [7].
We chose Malay as the worked example because a public, culturally grounded safety benchmark exists for it at a size small enough for the measurement problems to be visible and quantifiable [2, 7], and because national guidance on AI governance is in place in Malaysia [19]. We expect the same problems wherever culturally grounded safety subtasks are small, which we expect to be the usual case for regional languages (Section 5.3).
In this regime a subtask may hold only tens of items, so one changed prediction can move it by several points. Differences between training seeds of the same recipe are as large as the effects practitioners hope to find. Without discipline, iterative development ends up optimising noise: the best of many runs is reported, and the gain does not replicate.
This paper reports what the Malay safety track of SEA-HELM [1, 2] can support at its current size and published precision. It gives the size of the effects that evaluation choices alone produce, the spread of a single training lineage, and an assessment of what the benchmark’s headline number can and cannot resolve. Training methods, data sources and model internals are out of scope. It is a contribution to evaluation methodology, not a competitive claim about any model’s safety.
Our contributions are:
SEA-HELM [1] is a public evaluation suite for Southeast Asian languages. The Malay safety competency of its leaderboard [2], one of the task groups the leaderboard reports, is what this paper calls the safety track. It averages two parts. The first is toxicity detection: 1,000 texts, 500 toxic and 500 non-toxic. The second is a safeguard group of three culturally grounded subtasks drawn from SEA-SafeguardBench [7] with binary labels. Each asks whether a prompt or a response is harmful. Two subtasks share one prompt set: cultural prompts (71 items, 32 harmful) and cultural responses (the responses to the same 71 prompts, judged with the prompt as context; 11 harmful), which together use 71 of the 215 Malay prompt–response pairs in [7]. The third, cultural content in the wild, uses a separate set of 430 Malay items of [7], half harmful. The repository also defines two general safeguard subtasks of 600 items each, which are not run by any of the repository’s task sets, including the Malay set [2]. The official scorer computes each subtask’s balanced accuracy, the mean of the per-class recalls, and rescales it against chance as (balanced accuracy − 1/k)/(1 − 1/k), floored at zero, where k is the number of classes. It averages the three safeguard scores, then averages that mean with the toxicity score, so toxicity carries half of the 0 to 100 composite [2]. Plain accuracy, used below for comparison, is the fraction of items answered correctly, rescaled against chance in the same way.
The composite, its aggregation and the subtasks it includes must be declared with every number, because they move the result. One checkpoint scored anywhere from 45.17 to 50.47 depending on the aggregation (Table 1).
Table 1: One checkpoint under four aggregations (composite, points).
| Aggregation of one checkpoint | Composite (points) |
|---|---|
| Class-balanced, chance-normalised; toxicity + 3 cultural safeguard subtasks (official) | 45.76 |
| Class-balanced, chance-normalised; toxicity + all 5 safeguard subtasks | 45.17 |
| Plain accuracy, chance-normalised; toxicity + 3 cultural safeguard subtasks | 50.47 |
| Plain accuracy, chance-normalised; toxicity + all 5 safeguard subtasks | 48.35 |
The 5.3-point spread is larger than the 3.3-point range of the 12 training variants in Section 3.3. Choosing an aggregation after seeing results is a form of selection. Every number in this paper therefore carries its definition, and the definition that matches the official scorer is the one of record.
The evaluation choices in Sections 3.1 and 3.2 change nothing about the weights. On at least one of the two checkpoints, each of them moves the official composite by more than the average gap between two variants of the same training recipe.
Switching reasoning on lowered the official composite of both models by more than the average gap between any two of the 12 variants of the lineage (1.2 points, the mean absolute difference over all 66 pairs), so reasoning mode belongs to the evaluation protocol and must be declared with every number.
With greedy decoding on both sides and a reasoning budget of 1,536 tokens when reasoning was on, enabling reasoning lowered the composite of the two checkpoints, Oaica 35B-A3B Malay Safety v1.0 260923 and Oaica 35B-A3B Malay v1.0 260923, by 1.5 and 2.3 points (paired 95% intervals for the change, reasoning on minus off: −5.5 to +3.4 and −7.4 to +3.4; Figure 1). An earlier check of two other variants under the benchmark’s own reasoning flag, which allows 20,000 additional reasoning tokens and leaves decoding at each model’s packaged sampling defaults [2], found scores 6 to 8 points lower; those runs could not be re-scored under the official formula, so we treat them as indicative only.
Individually neither drop is statistically significant at α = 0.05, since both paired intervals include zero, and the two checkpoints share a base model, so they are not independent evidence. Both point the same way, and each exceeds the 1.2-point average gap between variants given above.
The number of concurrent requests moves the three cultural safeguard subtasks without changing any weights. Re-running those subtasks with one concurrent request instead of four changed only four of the 568 predictions that could be matched between the two runs (0.7 per cent), yet moved their mean score by 3.56 points (69.47 to 73.03), worth about 1.8 points on the composite. For Oaica 35B-A3B Malay v1.0 260923, the same three subtasks moved the other way, about −0.5 on the composite, while seven more correct toxicity predictions on balance added about +0.7, a net change of about 0.2 points. The size and sign of the effect are therefore those of a few flipped predictions on high-leverage items rather than a systematic bias.
The 12 training variants of one lineage spanned 3.3 points on the official composite (mean 45.36, SD 1.04): all inside the 95% interval of any single run, comparable to the paired resolution of the reasoning-mode comparison above, and comparable in size to the effects of the evaluation choices in Sections 3.1 and 3.2. Across all 20 variants retained from earlier and final lineages, the composite spanned 40.8 to 47.2 points.
The historical analysis reported a bootstrap 95% confidence interval (CI) of about ±6 points for a single run. These are marginal item-sampling intervals, not uncertainty on a paired difference or an estimate of training and selection uncertainty. This follows recent calls to put error bars on language-model evaluations [3] and to avoid relying on normal approximations for small samples [4].
The safety-tuned checkpoint, Oaica 35B-A3B Malay Safety v1.0 260923, scored 47.19 on the official safety composite (95% CI 40.59 to 52.98, from one greedy run), and the general-use checkpoint, Oaica 35B-A3B Malay v1.0 260923, scored 45.64 (CI 39.22 to 52.20). Re-run through a production serving stack with single requests, they scored 47.11 and 44.94. Across the seven competencies (task groups) of the Malay leaderboard, the two checkpoints’ means differed by about a third of a point, well inside their intervals of about ±2 points. Our reproduction is not the leaderboard’s runner and does not reproduce every published entry, so neither the safety composites above nor these means can be compared like-for-like with the published scores in Section 5, and comparing their marginal intervals cannot determine pairwise significance. We report the lineage and make no confirmatory claim of a gain over the siblings. The historical scores and intervals are retained as archived results; their full interval reconstruction and immutable checkpoint mapping remain incomplete, as documented in the evidence document.
Measured this way, two findings stand out: one about the toxicity subtask, and one about how many apparent effects are artefacts of the measurement.
Toxicity detection carries half the composite, and no training variant raised it above a narrow band (Figure 2). Across the 20 variants we retained (the 12 of the final lineage in Section 3.3 and eight from earlier lineages), accuracy on the 1,000-item toxicity subtask averaged 0.600 (SD 0.008, range 0.577 to 0.614). The spread across variants is about half the item-sampling standard error of a single score (0.015). Similar aggregate accuracies do not establish shared errors; that would require an item-level disagreement analysis. Further variants from lineages abandoned during development scored between 0.50 and 0.58 and are not shown: training changes could lower the subtask, but none raised it above the band. Because the subtask is balanced and carries half the composite, each 0.01 of toxicity accuracy is worth one composite point.
Several results bear on why the scores stop there. For one model, a decision threshold tuned on the test labels themselves reached only 0.617 accuracy, against 0.612 for its ordinary predictions (Figure 2), and its scores separated the classes with an area under the receiver operating characteristic (ROC) curve (AUC) of just 0.632. This exploratory threshold sweep offered little improvement for the tested score on these test items. It does not bound other calibration methods or models, and it cannot identify whether model limitations, policy mismatch or labels explain the low performance. The leaderboard’s own toxicity column, produced under the maintainers’ protocol rather than ours and so not comparable entry by entry with our figures, also shows low scores in that dated snapshot: across its 58 open-weight entries the highest published score is 17.51 normalised points, about 0.59 balanced accuracy, and the median entry scores about 0.55 [2]. Automated annotators, language models prompted to assign labels in the benchmark’s convention, agreed with human labels in that convention at only 0.51 to 0.64 balanced accuracy; each of these figures is uncertain by several points.
Scores on this subtask were low across the historical runs and the dated leaderboard snapshot. These observations do not distinguish model limitations, annotation-policy mismatch, missing context or label error, and do not quantify their relative contributions. Automated annotators are not an independent human reference. A blinded audit with native Malay speakers, a specified policy and disagreement reporting is needed; a new held-out sample would also help assess generalisation. Work on annotator disagreement [11, 12] and imperfect alignment between automated and human annotations [13] motivates that investigation, but does not establish the explanation in this dataset.
Two cases show the size of such effects. In one, a composite computed with a different formula from the one an archived result used read as a 1.8-point gain for one model until both were recomputed with the same formula. In another, a result file with no record of its configuration read 47.27 against 45.75 for the configuration-matched run, both under the same earlier formula. Several effects in this paper happen to be about 1.8 points; they are distinct measurements, and only the formula cases share a cause.
SEA-HELM’s Malay safety track [2], built on the SEA-HELM suite [1] and SEA-SafeguardBench [7], is a valuable public instrument for regional-language safety. Its small classes make some scores sensitive to individual outcomes. Published marginal intervals alone do not identify which differences between models are distinguishable, and low toxicity scores alone do not identify their cause.
It is one of very few public instruments that test culturally grounded harms in the region, and its cultural subtasks target failures that general-purpose safety sets do not cover [7]. The weaknesses listed in Table 2 concern what its headline number can support, not whether it is worth running. Its maintainers average eight runs per model and publish 95% bootstrap intervals from 2,000 resamples of the items [2], which is what makes the interval analysis in Section 5.1 possible.
Table 2: Properties of the safety track that limit what its headline number can support.
| Property | What we observed | What it means for a reader |
|---|---|---|
| Subtask size | Two subtasks hold only 71 items each, and one of their classes only 11; the source benchmark reports its lowest annotator agreement for response labels [7] | One changed prediction moves such a subtask by up to 9 points and the composite by up to 1.5 points (Figure 3), so pairwise comparisons should account for which items changed |
| Shared prompts | Two cultural subtasks score the same 71 prompts, once at prompt level and once at response level | They evaluate different targets on shared prompts; uncertainty estimation should account for that dependence |
| Aggregation | The Malay task set scores three of the five safeguard subtasks defined in the repository; the formula and task set are documented there, not beside entries | A score quoted without its formula and task set cannot be compared (Table 1) |
| Toxicity labels | A binary subtask carries half the composite; no variant we ran, and no entry in the dated leaderboard snapshot under the maintainers’ own protocol, exceeds about 0.61 accuracy, and automated annotators do not reproduce the convention either | Model limitations, policy mismatch, missing context and label error remain possible explanations |
| Opposing subtasks | Across 51 archived runs (variants and intermediate checkpoints), cultural-prompt and in-the-wild scores were negatively correlated (r ≈ −0.66), alongside differences in how often a model predicts “harmful”; this association does not establish causation | Inspect per-class errors and decision policies before interpreting the association as a causal trade-off |
| Reasoning-mode protocol | The maintainers publish their generation policy (each model’s packaged defaults, otherwise the serving engine’s) and one reasoning budget, and flag reasoning models, but not the mode and decoding values applied to each entry [2] | Entries may be scored in modes their developers never tuned |
| Data availability | The Malay toxicity data are served from a gated dataset named in the toxicity task’s configuration file (aisingapore/Safety-Toxicity-Detection) [2]; its public card lists no Malay split, although its file listing has a Malay folder, and on 30 September 2026 the data could not be downloaded without approved access | Independent reproduction now depends on being granted access |
| Confidence intervals | The ten highest of the 58 entries in the leaderboard page’s data share a 6.6-point interval range (Figure 4), and the five highest in its default view a 5.5-point range | Marginal intervals alone do not determine paired significance |
Under the official scorer, away from the zero floor, one corrected prediction moves the composite by 50/(3n) points on a safeguard subtask and by 50/n points on toxicity, where n is the size of the item’s class within its subtask (Figure 3). One of the 11 harmful items in the cultural-response subtask therefore moves the composite by 1.5 points, 15 times as much as one toxicity item (0.10) and about 20 times as much as one in-the-wild item (0.078). The cultural-prompt and cultural-response subtasks score only 71 distinct prompts (142 of the 1,572 scored items, 9.0 per cent), yet this is where single predictions move the composite most. These figures are for a single run, which is how we evaluate; published entries average eight runs per item, so an answer that changes in one of the eight runs moves an entry by an eighth of these amounts.
The data behind the leaderboard’s Malay page hold 58 open-weight entries of up to 200 billion parameters (dated 18 September 2026), of which the page’s default view displays 18; we use the full set. Its ten highest safety scores, from 46.22 down to 42.47, have 95% intervals that all share the range 40.8 to 47.4 points (Figure 4). Non-overlap of individual intervals is a conservative test for a difference, so we also approximated a 95% interval for each pairwise difference from the published half‑widths, treating entries as independent: every approximate interval includes zero, and the largest standardised difference is 0.93. This is an independence sensitivity calculation, not a paired test: the models share evaluation items, and the necessary covariance is unavailable in the published marginal intervals. The five highest entries in the default view likewise share a range (40.8 to 46.3); the same independence approximation includes zero for each pair. Further down that view some approximate intervals exclude zero, subject to the same covariance limitation. We therefore cannot determine from those marginal intervals which leading entries differ statistically. A paired test on shared items [5], with appropriate treatment of multiple comparisons, is needed.
Read the composite together with its uncertainty and subtask results. A small observed gap is inconclusive without an appropriate comparison; it is not automatically noise or evidence of equivalence. For models scored on shared items, report a paired interval on the difference. Read toxicity separately and validate the annotation policy against the intended use.
Practitioners should optimise behaviour on development data that reflects their own policy, use the benchmark as an external check, and report per-subtask results with intervals.
The paper’s own results are for one language, Malay; the single cross-language probe in Section 5.4 is a transfer check under its own serving configuration, not a multi-language safety evaluation, and no ordering between languages is claimed. What transfers is the structure of the problem. Item leverage is arithmetic: under a balanced-accuracy scorer, one changed prediction moves a subtask by an amount set by the size of the item’s class, so any language whose culturally grounded safety subtasks hold tens of items has the same exposure, and the expression in Section 5.1 can be applied to its subtask sizes directly. The benchmark family we assessed spans several Southeast Asian languages [1, 7], and comparable culturally grounded safety sets now exist for Indonesian languages [9].
Three reporting practices, items 1 to 3 of the checklist in Section 5.5, carry over unchanged: state which scorer and task set produced a score, declare the reasoning mode and request concurrency, and report an interval beside every score. What does not transfer is magnitude. How large each effect is depends on a language’s subtask sizes and annotation convention, and has to be measured language by language. For the same reason, our reading of the toxicity labels (Section 4.1) is a hypothesis to test in other languages, not a result about them.
The leverage arithmetic in Section 5.1 is a property of the scorer, not of Malay, so it should apply to any language whose culturally grounded safety subtasks are small. To test that expectation rather than assert it, we ran the same harness on a second regional language, Indonesian, which the model was not tuned for. No Indonesian-targeted text entered the fine-tuning mix, so the figures below are a cross-language transfer probe, not a within-language result, and no Indonesian-specific tuning or evaluation-driven development was performed.
Both checkpoints of one lineage were measured under a single shared
configuration: Oaica 35B-A3B Malay Safety v1.0 260923 and the
open-weight base Qwen3.6-35B-A3B that it is fine-tuned from. The harness
was SEA-HELM v1.3.0 on 20 of the 21 Indonesian tasks (the gated
syntax-criteria task could not be obtained for either
model), served as Q8_0 quantised weights under llama.cpp rather than the
vLLM bf16 stack of the Malay leaderboard, with sampling on and reasoning
off, over eight runs each. The serving stack and the run count therefore
differ from the Malay runs elsewhere in this paper, and these scores are
not comparable entry by entry with them.
Table 3: Indonesian SEA-HELM competency scores for two checkpoints of one lineage, under one shared configuration (SEA-HELM v1.3.0, 20 of 21 tasks, Q8_0 under llama.cpp, sampling on, reasoning off): mean and standard deviation (SD) over eight runs. The composite is the official safety score (Section 2); the overall figure is the equal-weight mean of the nine competencies above it. The two columns are not ordered (Section 5.2).
| Competency | Oaica 35B-A3B Malay Safety v1.0 260923 | Qwen3.6-35B-A3B |
|---|---|---|
| Multi-turn (LLM-judged) | 63.33 ± 2.45 | 84.95 ± 0.85 |
| Natural language understanding (NLU) | 77.25 ± 1.30 | 73.96 ± 4.16 |
| Safety (composite) | 57.16 ± 0.62 | 50.47 ± 9.46 |
| Natural language generation (NLG) | 55.10 ± 0.14 | 54.67 ± 0.19 |
| Natural language reasoning (NLR) | 87.45 ± 1.13 | 81.06 ± 5.87 |
| Linguistic diagnostics | 54.96 ± 6.58 | 44.05 ± 22.75 |
| Instruction following | 85.48 ± 1.63 | 87.62 ± 3.09 |
| Knowledge | 72.50 ± 5.07 | 67.92 ± 9.79 |
| Cultural | 76.80 ± 1.01 | 57.51 ± 14.95 |
| Overall (nine-competency mean) | 70.00 ± 1.59 | 66.91 ± 6.83 |
Two observations echo the Malay findings. First, the interval widths follow the item-leverage pattern of Section 5.1: the multi-turn and NLG subtasks are tightly determined (SD near one point), while linguistic diagnostics, knowledge and the cultural subtask carry SDs of five to about twenty-three points, the wide spread being what small item classes produce. Second, the composite again sits in the middle of its range, with the toxicity subtask low on both checkpoints, consistent with the hypothesis of Section 4.1 and with its cause left unresolved here as there.
Two of the eight base-model runs are anomalous: they collapse together across every classification subtask (toxicity accuracy 0.50 against about 0.71 elsewhere), which inflates that column’s SD. Over the six consistent runs the overall mean is 70.48 (SD 3.32), close to the fine-tuned checkpoint’s 70.00 (SD 1.59). The clearest and most reproducible separation is in the opposite direction, on the judge-rated multi-turn task, and there the two intervals do not overlap in this configuration. We report that gap rather than claim it as an advantage: the probe was not designed as a comparison, the two checkpoints are a fine-tune and its own base rather than independent entries, and no paired interval on the difference was computed (Section 5.1), so no ordering between the checkpoints is claimed. Both series are reported only to show that the measurement behaves in a second language as Section 5.1 predicts.
The findings above reduce to a short checklist for reporting any safety score, in any language. It describes what to publish beside a number, not how a model is built.
The checklist is implemented as a small open-source tool,
benchlint (github.com/sprapp-com/benchlint).
It lints a results report for the first four points above, flagging a
missing interval or sample size, ranking language with no test behind
it, false precision and single-run claims, and it computes item leverage
and interval estimates from item-level outcomes. It is language- and
model-agnostic, depends only on the Python standard library, and is
released under Apache-2.0.
Sections 3 to 5 concern what the benchmark can support. The two checkpoints, Oaica 35B-A3B Malay Safety v1.0 260923 (safety-tuned) and Oaica 35B-A3B Malay v1.0 260923 (general use), are sparse mixture-of-experts models built on Qwen3.6-35B-A3B, with about 35 billion parameters in total and about 3 billion active for each token. Tables 4 to 7 and 9 and 10 report point estimates; the 95% intervals for the checkpoints’ official composites are given in Section 3.3, the MalayMMLU sample intervals in Tables 9 and 10, and the public entries’ intervals with the leaderboard [2]. The recomputed Table 7 panel has its own intervals. Other rows are descriptive point estimates without reproduced intervals in this release. The point estimates should be read with those intervals in mind. This section places the two checkpoints on the wider Malay, English and Chinese benchmark set and lists possible supervised pilot uses and the validation still required. Published leaderboard rows and our historical proxy measurements use different instruments; their numerical gaps are not comparative effect estimates (Section 8).
Table 4 places our historical measurements beside published leaderboard values as separate descriptive panels. Our smaller item sets, automated judge and cultural proxy differ from the official instruments. Their numerical gaps are not estimates of comparative capability, and neither interval overlap nor an interval excluding zero removes that mismatch. Within our own historical panel, the two checkpoints differ by about a third of a point on the mean; no paired interval for that difference is reported. Multi-turn dialogue, knowledge-intensive question answering and longer-context serving remain areas for further evaluation.
Table 4: Descriptive results from historical measurements and published leaderboard entries. The protocols and, for several competencies, instruments differ; numerical gaps are not comparative effect estimates. Our historical runs used reasoning off and smaller item sets. Published values were captured on 24 September 2026 [2]. The safety row uses the official composite on both sides. Cultural scores in our panel use a proxy task, not the leaderboard’s cultural subtasks. The SEA-LION entry is the Qwen-based v4.5 27B release listed on the leaderboard [2]; Table 10 names different, Gemma-based releases.
| Competency / instrument | Historical Oaica 35B-A3B Malay Safety v1.0 260923 | Historical Oaica 35B-A3B Malay v1.0 260923 | Published GLM 5.1 | Published SEA-LION v4.5 27B |
|---|---|---|---|---|
| Malay mean (7 competencies) | 71.1 | 71.4 | 75.2 | 75.2 |
| Cultural (our proxy / published official) | 76.8 | 76.9 | 69.8 | 67.6 |
| Safety (official composite) | 47.2 | 45.6 | 48.1 | 40.6 |
| Translation (NLG) | 91.0 | 90.8 | 91.5 | 92.4 |
| Natural-language understanding (NLU) | 67.2 | 68.2 | 71.1 | 72.3 |
| Instruction following | 79.1 | 78.1 | 85.8 | 88.3 |
| Knowledge | 71.3 | 72.7 | 76.1 | 83.3 |
| Multi-turn chat | 65.0 | 67.6 | 83.9 | 82.1 |
Table 5 describes the historical Malay measurements. On Malay knowledge multiple-choice, the checkpoints score 82.8 and 83.4 against 76.8 for the earlier-stage checkpoint they were built from. This is a separate instrument from the published MalayMMLU scores and the leaderboard’s Knowledge competency. Malay reading-comprehension scores are 88.3 and 89.9; Malay instruction following scores 78 to 79 on a 105-item set, with the highest retained variant at 85.7; English instruction following scores 93.3 to 94.3 on development checkpoints (Table 6). These are descriptive results on the listed tests, not broad capability or deployment claims.
The same lineage also includes Chinese measurements. Development checkpoints scored 93.3 for Chinese-to-Malay and 90.5 to 90.8 for Malay-to-Chinese on a 200-item automatic metric using our protocol; this does not establish parity with other translation evaluations. On Chinese knowledge multiple-choice, they scored 84.4 against 79.7 for the earlier-stage checkpoint (Table 6). English-to-Malay scores were 90.6 to 90.8 and Malay-to-English 91.4 to 92.3 on those checkpoints. The two released checkpoints have aggregate translation scores of 90.8 to 91.0.
Table 5: The wider Malay benchmark set for the two released checkpoints. Instruction following is scored on subsets of about 105 items; the toxicity row is balanced accuracy on the 1,000-item subtask, where the retained historical variants occupy the range reported in Section 4.1; all rows are 0 to 100 points except the toxicity row, which is 0 to 1.
| Instrument | Oaica 35B-A3B Malay Safety v1.0 260923 | Oaica 35B-A3B Malay v1.0 260923 | Note |
|---|---|---|---|
| MalayMMLU (Malay knowledge, multiple-choice questions) | 83.4 | 82.8 | Earlier-stage checkpoint 76.8; other published MalayMMLU scores use different protocols and are not reproduced here |
| Belebele (Malay reading comprehension) | 88.3 | 89.9 | Historical point estimates |
| Sentiment (3-class, Malay) | 46.1 | 46.5 | Scored with reading comprehension in the NLU column |
| IFEval (Malay instruction following) | 79.1 | 78.1 | Best variant 85.7 |
| Toxicity (Malay) | 0.60 | 0.61 | Balanced accuracy; Historical variants ≤ about 0.61 |
Table 6: Translation and Chinese-language results, and English instruction following. Rows are measured on development checkpoints of the same lineage unless the row states otherwise (the aggregate row is for the two released checkpoints); translation is MetricX-style automatic scoring, 200 items per direction.
| Instrument | Score | Note |
|---|---|---|
| Translation, Chinese to Malay | 93.3 | Own protocol; no cross-protocol parity claim |
| Translation, Malay to Chinese | 90.5 to 90.8 | |
| Translation, English to Malay | 90.6 to 90.8 | |
| Translation, Malay to English | 91.4 to 92.3 | |
| Translation, aggregate for the released checkpoints | 90.8 to 91.0 | The only translation figure reported for the two released checkpoints |
| CMMLU (Chinese knowledge, multiple-choice questions) | 84.4 | Earlier-stage checkpoint 79.7 |
| IFEval (English instruction following) | 93.3 to 94.3 |
The historical serving measurements used one NVIDIA A100-80GB, vLLM
0.30.0, bfloat16 or 4-bit weights, a 2,864-token prompt and up to 200
output tokens. With one concurrent stream and prefix-cache hits
excluded, effective prefill throughput (prompt tokens divided by time to
first content token) was 14,300 to 15,600 tokens/s, with a first token
after about 0.2 s; single-stream decoding was 152 to 158 tokens/s. At
eight concurrent streams, aggregate decoding was 347 to 544 tokens/s on
the long prompt and 572 to 917 on an 89-token prompt. These
workload-specific observations are not latency guarantees, and they are
historical measurements on the previous serving stack. Our current
.oqm serving engine targets Blackwell-class hardware
(sm_120, RTX PRO 6000), with FP8-only kernels and split-KV and
shared-quantisation paths that replace the A100 baseline; the A100
figures are retained only as a historical reference point. The 4-bit
format reduced weight memory by about 3.2 times; historical MalayMMLU
point estimates changed from 83.2 to 82.6 and 83.0 to 82.4 on the
serving sample. No equivalence test is reported. Weight fit does not
include every context-length or concurrency requirement: runtime and
KV-cache memory must also be budgeted. Fully local operation requires a
deployment configured without external inference, telemetry or other
outbound integrations. The same weights are not tied to one accelerator
class: our .oqm serving engine can pin the model’s
mixture-of-experts layers in host CPU RAM, filling the available GPU
memory automatically and offloading the remainder, so the family also
runs on laptop- and desktop-class GPUs at reduced context and scales to
datacenter servers. The weights ship in our own .oqm
quantised container, and higher-throughput engine configurations are
offered for API and cloud inference; their throughput figures are
outside these measurements.
Reasoning off is a practical configuration to evaluate for short classification outputs and latency, but the paired reasoning-mode intervals in Section 3.1 include zero, so this study does not establish a safety advantage. Historical served safety point estimates were 48.9 and 44.9 with 4-bit weights, compared with 47.1 and 44.9 in bfloat16. These single-run results are not an equivalence or non-inferiority test. Historical benign over-refusal rates ranged from 1.7 to 8.5 per cent across variants; the range alone does not establish harmful-content recall, subgroup performance or deployment suitability.
For the archived runs, the reported safety composites are 47.2 and 45.6. Those runs differ from the recomputed panel in Table 7 and should not be merged into a single matched comparison. Holding a safeguard score S fixed, a composite above 50 requires toxicity accuracy above 1 − S/200 on this balanced binary task. This is a conditional arithmetic relationship, not evidence of a model or benchmark ceiling. Both components can change with a new model, protocol or sample.
Table 7 reports the saved runs recomputed from verified item files. The unsupported 41.8 stock, 42.2 peer and 45.3 specialist figures could not be tied to matching item outcomes in the available packet and are not retained as alternate estimates. The general-model run here is also distinct from the historical 45.6 result in Section 3.3. Product names follow the local run records; immutable checkpoint-to-product mapping has not yet been verified.
Table 7: Local safety runs, recomputed from verified item files with the SEA-HELM balanced-accuracy formula (toxicity plus three cultural subtasks). Greedy generation, reasoning off, batch size 8; 1,572 scored items per run. Intervals use 20,000 conditional item-bootstrap replicates, retaining the shared prompt/response structure, seed 20261004. They exclude training, generation-rerun and model-selection uncertainty. Model/dataset revisions and exact runtime provenance remain incomplete. These are not official leaderboard submissions. The evidence document identifies every run by content hash and provides recomputation code.
| Run | Recorded model | Composite | Conditional 95% interval |
|---|---|---|---|
| Run 1 | Oaica 35B-A3B Malay v1.0 260923 (general) | 44.92 | 37.88–51.58 |
| Run 2 | Oaica 35B-A3B Malay Coder v1.0 260923 | 45.91 | 39.40–51.54 |
| Run 3 | Oaica 35B-A3B Malay Researcher v1.0 260923 | 45.24 | 38.60–51.24 |
| Run 4 | Stock Qwen3.6-35B-A3B reference | 43.37 | 36.92–48.95 |
| Run 5 | SEA-LION v4.5 27B (Qwen-based) | 43.00 | 36.07–49.22 |
No paired model-difference interval is reported, and no comparative safety claim is made. The shared formula does not by itself make generation protocols or checkpoint histories equivalent. The historical smaller-model references in the abridged edition (36.9 for Qwen3.5-4B and 21.9 for a recorded Nanbeige 3B base release) come from earlier records and are outside this recomputed panel. Full-set first-token MalayMMLU reference measurements appear in Tables 9 and 10; they cannot estimate uplift over our historical generation-based sample scores.
Table 8 lists possible pilot applications. These are the company’s proposed uses, not validated deployment verdicts. Each requires evaluation on the customer’s policy, workload and affected communities, including harmful-content misses as well as false alarms. Memory fit and throughput establish technical feasibility only.
Table 8: Candidate uses and the validation required before deployment.
| Candidate use | Evidence available | Validation still required |
|---|---|---|
| Local Malay moderation and screening | Historical classification scores and historical A100 serving measurements | Per-class recall/false alarms, subgroup and code-switching errors, policy alignment, adversarial tests and human escalation |
| Chat guardrail | Short-output latency measured on specified prompts | Tail latency under load, harmful-content misses, bypasses and uncertain-case routing |
| Culturally sensitive content | A Malay natural-language-inference proxy | Native-speaker policy review and task-specific harmful/benign examples; the proxy is not a moderation validation |
| Malay, English and Chinese translation | Automatic scores on development checkpoints | Human adequacy checks and the actual served checkpoint, domain and language directions |
| Business and government drafting | Historical language and instruction-following tests | Factuality, source verification, confidentiality and human approval of consequential outputs |
| Local policy adaptation | Locally deployable weights and internal development capability | Contractual rights, auditable training lineage and fresh holdout evaluation after changes |
Published MalayMMLU [22] scores use different evaluation protocols and were not reproduced in this study. We therefore do not place vendor-reported results beside our measurements or infer relative performance from them. A meaningful comparison would require the same question set, prompt, answer extraction, model version and scoring procedure.
Oaica 35B-A3B Malay is fine-tuned from Qwen3.6-35B-A3B, the open-weight base, and the stock reference run measures that base directly. Which checkpoint produced each historical result, and the base’s exact revision tag, remain to be mapped. Under one full-set first-token protocol and template, recorded outside this release and not tabulated here, the stock model scores 80.7 on MalayMMLU, the smaller Qwen3.5-4B scores 71.1 and Nanbeige4.2-3B scores 51.5. The sizes differ, so this is not a ranking of the families; it shows how far apart stock models of these sizes score under one protocol, under that evaluation method. These stock figures were measured after our checkpoints were built, so they are retrospective reference measurements, not a record of the original model-selection decision.
The historical Oaica figures of 83.4 and 82.8 use generation with reasoning on and a 500-question sample; stock 80.7 uses a full-set first-token protocol. No matched MalayMMLU uplift estimate follows from those numbers. Against an earlier-stage checkpoint, the historical record reports changes from 76.8 to 82.8 and 83.4 on Malay knowledge, and from 79.7 to 84.4 on Chinese knowledge on a development checkpoint. These are protocol-specific point differences without paired intervals here. The recomputed safety runs in Table 7 also establish no comparative deployment-safety claim or causal effect of our training recipe.
Question. Does an experimental Malay fine-tune of Qwen3.5-2B, which we call Oaica 2B Malay, score higher on MalayMMLU than its stock base, and does the answer depend on the evaluation protocol? The model is not a release candidate, and we report no safety composite for it.
Setup. Both models ran zero-shot on one RTX 3090 with one inference engine (vLLM) in bfloat16, with greedy decoding and the same Malay instruction prompt in every protocol. The Qwen3.5-2B chat template, shared by both models, emits an empty reasoning block unless reasoning is explicitly enabled, unlike the larger Qwen models we evaluate, so every chat-template run here is reasoning off; we did not run either model with reasoning on. We used three protocols:
Table 9: MalayMMLU for the stock Qwen3.5-2B and Oaica 2B Malay under three protocols (points, 0 to 100). Each G figure has a 95% interval of up to about ±4.4 points and each first-token figure about ±0.6 points; we have not computed the paired interval on a difference. F-chat uses first-token answer-letter selection with a no-think chat template; G is our own protocol and is not comparable with Table 5. Differences are computed from unrounded values.
| Protocol | Items | Stock Qwen3.5-2B | Oaica 2B Malay | Difference |
|---|---|---|---|---|
| G, generation with answer extraction | 500 | 56.0 | 62.4 | +6.4 |
| F-plain, first token, plain prompt | 24,213 | 42.4 | 52.0 | +9.6 |
| F-chat, first token, chat template | 24,213 | 59.3 | 44.1 | −15.3 |
Findings.
1. The sign of the comparison depends on the protocol. The tuned model scores higher under G and F-plain and lower under F-chat. The two first-token differences are the best resolved, each on 24,213 items per model; the G difference, on 500 items, is the least resolved of the three.
2. The protocol alone moves a single model by more than any difference between the two models in Table 9. The stock model ranges over 17.0 points across the three protocols (ranges are computed from unrounded values), between the two full-set protocols and therefore well resolved, against between-model differences of 6.4, 9.6 and 15.3 points in absolute value. The tuned model ranges over 18.3 points, which depends on the 500-question G figure and is indicative; between the two full-set protocols it moves by 7.9 points.
3. Extraction failures do not explain the generation result. No letter could be extracted from 0.6 per cent of the stock model’s outputs (three items; scoring them all correct would raise the stock figure to 56.6) and from none of the tuned model’s.
Across other models. We applied the same three protocols to five further public models on a second host of the same GPU type (Table 10; the Qwen3.5-2B and Oaica 2B rows repeat Table 9).
Table 10: MalayMMLU under the three protocols of Table 9 for seven models, ordered by stored parameter count; the order is not a ranking. G is generation with answer extraction on the 500-question sample, with the chat template’s default setting and a 1,536-token budget (up to about ±4.4 points per figure); F-plain and F-chat are first-token scores on all 24,213 questions (about ±0.6 points), and F-chat is reasoning off where the template allows it. F-plain omits the chat template, so its figures are not the protocol under which these instruction-tuned models are meant to be used and should not be quoted as a model’s MalayMMLU score. ¹ For Qwen3-1.7B, G is reasoning on (its template default) and F-chat is reasoning off (53.5; 44.8 under its default template), so its range includes that switch; with reasoning off in generation too (57.2, 64-token budget) the range is 10.8. ² The 4.1 release of Nanbeige is the release tabulated here; the 4.2 post-trained release appears only in the historical first-token reference figures quoted in Section 6.7. It begins every answer with a reasoning block of its own accord and its template has no switch to suppress it, so G is reasoning on, and its F-chat figure is the letter forced at the first position, where the model would otherwise open that block. Figures are our own measurements of public models and may differ from the developers’ reports. Range is the highest minus the lowest of the three columns, from unrounded values. Extraction failed on at most 1.6 per cent of any model’s G outputs. Points, 0 to 100.
| Model (stored parameters, billions) | G, generation | F-plain, first token, plain prompt | F-chat, first token, chat template | Range across protocols |
|---|---|---|---|---|
| Qwen3.5-0.8B (0.87) | 50.4 | 46.9 | 51.8 | 4.9 |
| Qwen3-1.7B¹ (2.03) | 63.6 | 46.4 | 53.5 | 17.2 |
| Oaica 2B Malay (2.21) | 62.4 | 52.0 | 44.1 | 18.3 |
| Qwen3.5-2B (stock) (2.27) | 56.0 | 42.4 | 59.3 | 17.0 |
| Nanbeige4.1-3B² (3.93) | 44.0 | 45.4 | 36.9 | 8.5 |
| Gemma-SEA-LION-v4-4B-VL (4.30) | 64.2 | 64.7 | 65.8 | 1.6 |
| Gemma-SEA-LION-v4.5-E2B-IT (5.12) | 68.2 | 41.9 | 68.4 | 26.5 |
4. The sensitivity is not specific to our model. The range across the three protocols runs from 1.6 to 26.5 points. Five of the seven models move by more than 8 points, and four by at least 16.9 points.
5. The model with the highest figure depends on the protocol. The top figure under G and F-chat belongs to Gemma-SEA-LION-v4.5-E2B-IT and under F-plain to Gemma-SEA-LION-v4-4B-VL, where the E2B figure is among the two lowest (41.9, with the stock Qwen3.5-2B at 42.4; no paired test of that gap is reported); we draw no ranking from this.
6. G and F-chat agree within G’s interval for the two SEA-LION models, Qwen3.5-0.8B and the stock Qwen3.5-2B (gaps of 0.2 to 3.3 points), whereas the clear gaps are those of Oaica 2B Malay (18.3), Qwen3-1.7B (10.1, with G reasoning on) and Nanbeige4.1-3B (7.1). We did not investigate why.
Hypothesis, not tested. Fine-tuning may have changed how the model begins an answer. A first-token protocol under the chat template would penalise that, whereas a plain prompt that ends in the answer cue would not, which fits the direction of F-plain and F-chat. In the one stored sample output the tuned model opens with “Jawapan:” before the letter and the stock model with the letter, which is consistent with this reading but is a single example, not a test. A direct test, comparing the model’s unconstrained first token with its letter-restricted choice, has not been run.
What we conclude. We claim no improvement over the stock base: the evidence on this model is protocol-dependent, and one model pair and one sample cannot settle which protocol is closest to its real ability. The tuned-versus-earlier-stage gains of Section 6.7 were measured under one protocol only and should be read as specific to it. For future comparisons of a tuned model with its base we will report at least two protocols and treat a comparison whose sign changes between them as unresolved.
This work sits where five lines of research meet: statistics for language-model evaluation, culturally grounded safety for Southeast Asian languages, annotator disagreement, the safety of reasoning models, and the validity of benchmarks and leaderboards; we close with the Malaysian governance context.
Statistics for evaluations. Miller argues that evaluation results should carry error bars and be compared with paired analyses [3]. Bowyer et al. show that intervals based on the central limit theorem are unreliable on benchmarks with fewer than a few hundred items [4], and Wein et al. estimate significance for leaderboard differences from a single test set [5]. We apply these ideas to regional safety evaluation.
Regional and culturally grounded safety. SEA-HELM [1] evaluates Southeast Asian languages holistically; its leaderboard [2] has since added Malay, safeguard subtasks drawn from SEA-SafeguardBench [7], and bootstrap intervals over eight runs per model. Recent work builds culturally grounded safeguards and safety benchmarks for the region, including SEA-Guard [6], SEA-SafeguardBench [7], SEALGuard [8] and IndoSafety [9], and a joint testing exercise by several AI safety institutes covered Malay [10]. Our assessment complements this work by quantifying what the Malay safety track can resolve.
Annotator disagreement. Safety evaluation increasingly treats disagreement among annotators as information about genuine ambiguity rather than as noise [11], and recent modelling work separates such systematic disagreement from annotator error [12]. Automated safety annotations have also been compared with human ones, with imperfect agreement [13]. Our toxicity finding is consistent with this work.
Reasoning and safety. Studies of reasoning models examine the monitorability of reasoning traces and the safety of long chains of thought [14, 15]. We add an evaluation-side observation: switching reasoning on lowered the safety composite of both checkpoints (by 1.5 and 2.3 points, both historical paired intervals including zero), so the mode must be declared with every score.
Benchmark and leaderboard validity. Recent work asks what benchmarks measure and what their rankings support: Desai et al. examine validity across fifty-six AI benchmarks [16], and Yang and Chen study leaderboard claims under hidden model selection [17]. On SWE-bench Verified, Liu et al. report that paired tests do not distinguish any of the 29 adjacent pairs among the top thirty entries; on the larger Test split, some adjacent pairs are distinguishable [18]. Our assessment asks a related question of a regional-language safety track: which published differences exceed the benchmark’s own measurement uncertainty.
Governance context. Evaluation records can inform organisational governance and procurement. This study does not establish regulatory compliance or government endorsement. References [19, 20] provide the dated Malaysian policy context; legislative forecasts are not a basis for our product claims.
The evidence comes from one primary language with a single cross-language probe (Section 5.4), one benchmark family and, for the safety findings, one base model, under a compute budget that limited replication. Other model families appear only as published leaderboard entries (the leaderboard-wide toxicity and interval results of Sections 4.1 and 5.1, Figure 4, and the two entries beside ours in Table 4), as instrument reference points on the official scorer and on MalayMMLU (Sections 6.4, 6.6 and 6.7; Tables 7, 9 and 10) and in the non-safety protocol-sensitivity measurements of Section 6.8. Several findings therefore may not generalise.
Safety work handles harmful text and references to real communities, so the discipline constrains what is stored, published and released.
In this case study, some observed safety-score differences induced by evaluation settings were comparable to differences among closely related models. In the measurements reported here, aggregation moved one checkpoint by 5.3 points, request concurrency shifted the three cultural safeguard subtasks of one model by an amount worth about 1.8 composite points, and reasoning mode lowered the composite by 1.5 and 2.3 points, with paired intervals that include zero. By comparison, 12 training variants of one lineage spanned 3.3 points and differed pairwise by 1.2 points on average.
Toxicity performance remained low in the evaluated historical runs, but these results do not distinguish model limitations, annotation-policy mismatch, missing context or label error. The published intervals of leading entries overlap; without paired item-level outcomes we cannot determine which differences are statistically distinguishable. This case study supports reporting scorer definitions, inference settings, per-class errors and uncertainty. It does not establish comparative deployment safety, a universal benchmark ceiling or defects in the labels.
Early-stage Malay evaluation in this work used GlossoBench [21] (see Competing interests).
The author owns Oaica, which builds Malay language models, offers them for on-premises deployment by arrangement, and offers evaluation consulting. The study includes the company’s checkpoints, training variants, archived runs and serving measurements; Section 6 is the company’s own assessment of its products, including candidate pilot uses and experimental-model measurements. Oaica also releases GlossoBench [21], an open evaluation harness used for early-stage Malay evaluation in this work. It describes SEA-HELM as its closest peer and is published under the GitHub account sprapp-com of Sprapp, the group of which Oaica is a part. The assessment of SEA-HELM in Sections 4 and 5 should be read with these interests in mind.
Oaica is a safety-focused AI company. It builds Malay language models and works with teams that need regional-language safety evaluation they can defend: audits of how a benchmark score was produced and how much of it the benchmark can support, reliability reviews of safety leaderboards, and on-premises Malay moderation consultation. De-identified evaluation tables, with their intervals, and item-leverage worksheets behind this paper are available on request. This paper: https://research.oaica.com/malay-llm-safety-score/. A regional-language edition without the product sections is at https://research.oaica.com/regional-language-llm-safety-malay-indonesian/. Enquiries: [email protected] (general enquiries: [email protected]).