@misc{wong2026-regional-language-llm-safety-malay-indonesian,
author = {Wong, Sam},
title = {AI Safety Evaluation in a Regional Language: A Case Study with a Fine-Tuned Malay Model and an Indonesian Cross-Language Probe},
year = {2026},
howpublished = {Oaica Research},
url = {https://research.oaica.com/2026/10/regional-language-llm-safety-malay-indonesian/}
}A measurement study of the SEA-HELM Malay safety track, an Indonesian cross-language probe, and a checklist for any language
Oaica (oaica.com)
2 October 2026
A safety score depends on its evaluation protocol and annotation policy. We examine those dependencies in one regional-language case study. Using the Malay safety track of SEA-HELM (Southeast Asian Holistic Evaluation of Language Models), we show that a single changed prediction can move a subtask score by up to 9 points and the composite by 1.5. In this case study, selected evaluation settings produced score differences comparable in size to variation among closely related historical runs. Using two checkpoints of a Malay large language model fine-tuned from the open-weight base Qwen3.6-35B-A3B, and 20 training variants retained from one lineage’s final runs and several earlier lineages, we document what the track can and cannot resolve, including examples where measurement effects complicate interpretation.
The choice of aggregation moved one checkpoint by 5.3 points, a change in request concurrency shifted the three cultural safeguard subtasks of one model by an amount worth about 1.8 composite points, and switching reasoning on lowered the composite by 1.5 and 2.3 points, both with 95% intervals on the paired difference that include zero. By comparison, 12 training variants of one lineage spanned 3.3 points, any two of them differed by 1.2 points on average, and a single run’s 95% interval was about ±6 points. We then assess the benchmark itself: toxicity performance remained low in the evaluated runs, small class sizes give individual items high leverage, and overlapping published intervals do not establish which model differences are statistically distinguishable. These observations do not identify defects in the labels or establish a universal performance ceiling. The item-leverage arithmetic transfers across languages when the task structure and class sizes are supplied; the checklist is a proposal for other evaluations to adapt. A cross-language probe in Indonesian, a language the model was not tuned for, reproduces the same measurement behaviour (Section 5.4). Training methods are out of scope. A text-free evidence document records the available evidence and unresolved provenance.
Keywords: LLM safety evaluation; regional language safety; benchmark reliability; Malay (Bahasa Melayu); Indonesian (Bahasa Indonesia); cross-language transfer; multilingual and culturally grounded safety; Southeast Asian languages; SEA-HELM; SEA-SafeguardBench; confidence intervals; annotation disagreement; leaderboard ranking uncertainty.
It does not rank models, and it does not claim that any model is safe, is the safest, or is safer than another. It does not describe training methods, data sources or model internals. Its evidence comes from one primary language, Malay, with a single cross-language probe in Indonesian (Section 5.4), one benchmark family and, apart from two reference points on other open models in Section 4.3, from one base model, Qwen3.6-35B-A3B (Section 7).
This paper assesses what a single published number can support, not the intent or quality of the work behind it. SEA-HELM and SEA-SafeguardBench are among very few public instruments for culturally grounded safety in the region, and the leaderboard maintainers’ practice of publishing intervals over repeated runs [2] is what made the interval analysis in Section 5 possible. We offer our evaluation tables and item-leverage worksheets on request.
Evidence notice. This is an observational case study, not an independent model certification. Historical scores are distinguished from the recomputed runs, which are reported in the named edition and the evidence document, not here. The evidence document supplies text-free counts, exact score-recomputation code and an explicit list of missing provenance. Historical interval reconstruction and the mapping of saved results to checkpoint hashes remain open; no claim of complete reproducibility or verified absence of contamination is made. This is an abridged edition of the same study as the Malay edition, not a separate experiment or independent replication.
In our experience with one regional language, safety work failed more often at measurement than at training. The harms that matter are culturally specific, and they concern religion, ethnicity, institutions and local politics. The public benchmarks that measure them are small: in Malay, two of the three cultural safeguard subtasks hold only 71 items each [2, 7]. Their labels necessarily encode one annotation convention, which other competent annotators may not share [7]. SEA-HELM and SEA-SafeguardBench are valuable public instruments [1, 7]; the scorer comparison in Section 2 and the interval analysis in Section 5 are possible because the leaderboard’s maintainers document the scorer and publish intervals over repeated runs [2], and the assessment in Section 5 draws on the annotator agreement that the source benchmark reports [7].
We use Malay as the worked example because a public, culturally grounded safety benchmark exists for it at a size small enough for the measurement problems to be visible and quantifiable [2, 7]. We expect the same problems wherever culturally grounded safety subtasks are small, which we expect to be the usual case for regional languages (Section 5.3).
In this regime a subtask may hold only tens of items, so one changed prediction can move it by several points. Differences between training seeds of the same recipe are as large as the effects practitioners hope to find. Without discipline, iterative development ends up optimising noise: the best of many runs is reported, and the gain does not replicate.
This paper reports what the Malay safety track of SEA-HELM [1, 2] can support at its current size and published precision, and gives the arithmetic and checklist that let a reader run the same checks on another language. It gives the size of the effects that evaluation choices alone produce, the spread of a single training lineage, and an assessment of what the benchmark’s headline number can and cannot resolve. Training methods, data sources and model internals are out of scope. It is a contribution to evaluation methodology, not a competitive claim about any model’s safety.
Our contributions are:
SEA-HELM [1] is a public evaluation suite for Southeast Asian languages. The Malay safety competency of its leaderboard [2], one of the task groups the leaderboard reports, is what this paper calls the safety track. It averages two parts. The first is toxicity detection: 1,000 texts, 500 toxic and 500 non-toxic. The second is a safeguard group of three culturally grounded subtasks drawn from SEA-SafeguardBench [7] with binary labels. Each asks whether a prompt or a response is harmful. Two subtasks share one prompt set: cultural prompts (71 items, 32 harmful) and cultural responses (the responses to the same 71 prompts, judged with the prompt as context; 11 harmful), which together use 71 of the 215 Malay prompt–response pairs in [7]. The third, cultural content in the wild, uses a separate set of 430 Malay items of [7], half harmful. The repository also defines two general safeguard subtasks of 600 items each, which are not run by any of the repository’s task sets, including the Malay set [2]. The official scorer computes each subtask’s balanced accuracy, the mean of the per-class recalls, and rescales it against chance as (balanced accuracy − 1/k)/(1 − 1/k), floored at zero, where k is the number of classes. It averages the three safeguard scores, then averages that mean with the toxicity score, so toxicity carries half of the 0 to 100 composite [2]. Plain accuracy, used below for comparison, is the fraction of items answered correctly, rescaled against chance in the same way.
The composite, its aggregation and the subtasks it includes must be declared with every number, because they move the result. One checkpoint scored anywhere from 45.17 to 50.47 depending on the aggregation (Table 1).
Table 1: One checkpoint under four aggregations (composite, points).
| Aggregation of one checkpoint | Composite (points) |
|---|---|
| Class-balanced, chance-normalised; toxicity + 3 cultural safeguard subtasks (official) | 45.76 |
| Class-balanced, chance-normalised; toxicity + all 5 safeguard subtasks | 45.17 |
| Plain accuracy, chance-normalised; toxicity + 3 cultural safeguard subtasks | 50.47 |
| Plain accuracy, chance-normalised; toxicity + all 5 safeguard subtasks | 48.35 |
The 5.3-point spread is larger than the 3.3-point range of the 12 training variants in Section 3.3. Choosing an aggregation after seeing results is a form of selection. Every number in this paper therefore carries its definition, and the definition that matches the official scorer is the one of record.
The evaluation choices in Sections 3.1 and 3.2 change nothing about the weights. On at least one of the two checkpoints, each of them moves the official composite by more than the average gap between two variants of the same training recipe.
Switching reasoning on lowered the official composite of both models by more than the average gap between any two of the 12 variants of the lineage (1.2 points, the mean absolute difference over all 66 pairs), so reasoning mode belongs to the evaluation protocol and must be declared with every number.
With greedy decoding on both sides and a reasoning budget of 1,536 tokens when reasoning was on, enabling reasoning lowered the composite of the two checkpoints, Oaica 35B-A3B Malay Safety v1.0 260923 and Oaica 35B-A3B Malay v1.0 260923, by 1.5 and 2.3 points (paired 95% intervals for the change, reasoning on minus off: −5.5 to +3.4 and −7.4 to +3.4; Figure 1). An earlier check of two other variants under the benchmark’s own reasoning flag, which allows 20,000 additional reasoning tokens and leaves decoding at each model’s packaged sampling defaults [2], found scores 6 to 8 points lower; those runs could not be re-scored under the official formula, so we treat them as indicative only.
Individually neither drop is statistically significant at α = 0.05, since both paired intervals include zero, and the two checkpoints are fine-tuned from the same base, Qwen3.6-35B-A3B, so they are not independent evidence. Both point the same way, and each exceeds the 1.2-point average gap between variants given above.
The number of concurrent requests moves the three cultural safeguard subtasks without changing any weights. Re-running those subtasks on the production serving stack of Section 3.3 with one concurrent request instead of four changed only four of the 568 predictions that could be matched between the two runs (0.7 per cent), yet moved their mean score by 3.56 points (69.47 to 73.03), worth about 1.8 points on the composite. For Oaica 35B-A3B Malay v1.0 260923, the same three subtasks moved the other way, about −0.5 on the composite, while seven more correct toxicity predictions on balance added about +0.7, a net change of about 0.2 points. The size and sign of the effect are therefore those of a few flipped predictions on high-leverage items rather than a systematic bias.
The 12 training variants of one lineage spanned 3.3 points on the official composite (mean 45.36, SD 1.04): all inside the 95% interval of any single run, comparable to the paired resolution of the reasoning-mode comparison above, and comparable in size to the effects of the evaluation choices in Sections 3.1 and 3.2. Across all 20 variants retained from earlier and final lineages, the composite spanned 40.8 to 47.2 points.
The historical analysis reported a bootstrap 95% confidence interval (CI) of about ±6 points for a single run. These are marginal item-sampling intervals, not uncertainty on a paired difference or an estimate of training and selection uncertainty. This follows recent calls to put error bars on language-model evaluations [3] and to avoid relying on normal approximations for small samples [4].
The safety-tuned checkpoint, Oaica 35B-A3B Malay Safety v1.0 260923, scored 47.19 on the official safety composite (95% CI 40.59 to 52.98, from one greedy run), and the general-use checkpoint, Oaica 35B-A3B Malay v1.0 260923, scored 45.64 (CI 39.22 to 52.20). Re-run through a production serving stack with single requests and 16-bit weights, they scored 47.11 and 44.94. Across the seven competencies (task groups) of the Malay leaderboard, the two checkpoints’ means differed by about a third of a point, well inside their intervals of about ±2 points. Our reproduction is not the leaderboard’s runner and does not reproduce every published entry, so neither the safety composites above nor these means can be compared like-for-like with the published scores in Section 5, and comparing their marginal intervals cannot determine pairwise significance. We report the lineage and make no confirmatory claim of a gain over the siblings. The historical scores and intervals are retained as archived results; their full interval reconstruction and immutable checkpoint mapping remain incomplete, as documented in the evidence document.
Measured this way, two findings stand out: one about the toxicity subtask, and one about how many apparent effects are artefacts of the measurement.
Toxicity detection carries half the composite, and no training variant raised it above a narrow band (Figure 2). Across the 20 variants we retained (the 12 of the final lineage in Section 3.3 and eight from earlier lineages), accuracy on the 1,000-item toxicity subtask averaged 0.600 (SD 0.008, range 0.577 to 0.614). The spread across variants is about half the item-sampling standard error of a single score (0.015). Similar aggregate accuracies do not establish shared errors; that would require an item-level disagreement analysis. Further variants from lineages abandoned during development scored between 0.50 and 0.58 and are not shown: training changes could lower the subtask, but none raised it above the band. Because the subtask is balanced and carries half the composite, each 0.01 of toxicity accuracy is worth one composite point.
Several results bear on why the scores stop there. For one model, a decision threshold tuned on the test labels themselves reached only 0.617 accuracy, against 0.612 for its ordinary predictions (Figure 2), and its scores separated the classes with an area under the receiver operating characteristic (ROC) curve (AUC) of just 0.632. This exploratory threshold sweep offered little improvement for the tested score on these test items. It does not bound other calibration methods or models, and it cannot identify whether model limitations, policy mismatch or labels explain the low performance. The leaderboard’s own toxicity column, produced under the maintainers’ protocol rather than ours and so not comparable entry by entry with our figures, also shows low scores in that dated snapshot: across its 58 open-weight entries the highest published score is 17.51 normalised points, about 0.59 balanced accuracy, and the median entry scores about 0.55 [2]. Automated annotators, language models prompted to assign labels in the benchmark’s convention, agreed with human labels in that convention at only 0.51 to 0.64 balanced accuracy; each of these figures is uncertain by several points.
Scores on this subtask were low across the historical runs and the dated leaderboard snapshot. These observations do not distinguish model limitations, annotation-policy mismatch, missing context or label error, and do not quantify their relative contributions. Automated annotators are not an independent human reference. A blinded audit with native Malay speakers, a specified policy and disagreement reporting is needed; a new held-out sample would also help assess generalisation. Work on annotator disagreement [11, 12] and imperfect alignment between automated and human annotations [13] motivates that investigation, but does not establish the explanation in this dataset.
Two cases show the size of such effects. In one, a composite computed with a different formula from the one an archived result used read as a 1.8-point gain for one model until both were recomputed with the same formula. In another, a result file with no record of its configuration read 47.27 against 45.75 for the configuration-matched run, both under the same earlier formula. Several effects in this paper happen to be about 1.8 points; they are distinct measurements, and only the formula cases share a cause.
We also ran two public open-weight models of 3 to 4 billion parameters through the same harness configuration and scorer (Table 2). These are historical reference measurements, not a ranking or like-for-like comparison: the Nanbeige base release is recorded as a base release, and the small set of runs cannot establish a general relationship between model size and toxicity performance.
Table 2: Historical official safety composites and component scores for two open-weight reference models (Qwen3.5-4B and a Nanbeige 3B base release) and Oaica 35B-A3B Malay Safety v1.0 260923 and Oaica 35B-A3B Malay v1.0 260923, all calculated with the same official scorer. The small-model measurements are descriptive references; the Nanbeige base release is recorded as a base release, and which checkpoint produced each historical result remains to be mapped. Because the toxicity gold set is balanced, raw accuracy equals balanced accuracy, and the toxicity half is 2 × accuracy − 1 rescaled to 100. Points, 0 to 100, except toxicity accuracy, which is 0 to 1.
| Model | Safeguard half | Toxicity half | Composite | Toxicity accuracy |
|---|---|---|---|---|
| Qwen3.5-4B (instruction-tuned, 4 billion parameters) | 68.5 | 5.2 | 36.9 | 0.526 |
| Nanbeige 3B base release (recorded, nominally 3 billion parameters) | 34.4 | 9.4 | 21.9 | 0.547 |
| Oaica 35B-A3B Malay Safety v1.0 260923 (safety-tuned checkpoint) | 74.6 | 19.8 | 47.2 | 0.599 |
| Oaica 35B-A3B Malay v1.0 260923 (general-use checkpoint) | 70.1 | 21.2 | 45.6 | 0.606 |
In these two runs the safeguard half was higher than the toxicity half. On the three cultural safeguard subtasks the instruction-tuned smaller model scores 68.5 and the base-release model 34.4; on normalised toxicity they score 5.2 and 9.4, and the composite is the mean of the two halves. Toxicity accuracy is close to chance for both, 0.526 and 0.547; because the gold labels are balanced, the toxicity half simply restates those accuracies, so both models sit below the band of Section 4.1 rather than above it, and the reading given there applies to them as well, with the same caveats. The historical runs tabulated here do not exceed about 0.61 balanced accuracy on this subtask. This is an observed range, not a ceiling.
Holding the safeguard half S fixed, a composite above 50 requires toxicity accuracy above 1 − S/200 on this balanced binary task. This conditional arithmetic does not bound future models or datasets: both halves can change. The historical observations establish no universal model or benchmark ceiling.
SEA-HELM’s Malay safety track [2], built on the SEA-HELM suite [1] and SEA-SafeguardBench [7], is a valuable public instrument for regional-language safety. Its small classes make some scores sensitive to individual outcomes. Published marginal intervals alone do not identify which differences between models are distinguishable, and low toxicity scores alone do not identify their cause.
It is one of very few public instruments that test culturally grounded harms in the region, and its cultural subtasks target failures that general-purpose safety sets do not cover [7]. The weaknesses listed in Table 3 concern what its headline number can support, not whether it is worth running. Its maintainers average eight runs per model and publish 95% bootstrap intervals from 2,000 resamples of the items [2], which is what makes the interval analysis in Section 5.1 possible.
Table 3: Properties of the safety track that limit what its headline number can support.
| Property | What we observed | What it means for a reader |
|---|---|---|
| Subtask size | Two subtasks hold only 71 items each, and one of their classes only 11; the source benchmark reports its lowest annotator agreement for response labels [7] | One changed prediction moves such a subtask by up to 9 points and the composite by up to 1.5 points (Figure 3), so pairwise comparisons should account for which items changed |
| Shared prompts | Two cultural subtasks score the same 71 prompts, once at prompt level and once at response level | They evaluate different targets on shared prompts; uncertainty estimation should account for that dependence |
| Aggregation | The Malay task set scores three of the five safeguard subtasks defined in the repository; the formula and task set are documented there, not beside entries | A score quoted without its formula and task set cannot be compared (Table 1) |
| Toxicity labels | A binary subtask carries half the composite; no variant we ran, and no entry in the dated leaderboard snapshot under the maintainers’ own protocol, exceeds about 0.61 accuracy, and automated annotators do not reproduce the convention either | Model limitations, policy mismatch, missing context and label error remain possible explanations |
| Opposing subtasks | Across 51 archived runs (variants and intermediate checkpoints), cultural-prompt and in-the-wild scores were negatively correlated (r ≈ −0.66), alongside differences in how often a model predicts “harmful”; this association does not establish causation | Inspect per-class errors and decision policies before interpreting the association as a causal trade-off |
| Reasoning-mode protocol | The maintainers publish their generation policy (each model’s packaged defaults, otherwise the serving engine’s) and one reasoning budget, and flag reasoning models, but not the mode and decoding values applied to each entry [2] | Entries may be scored in modes their developers never tuned |
| Data availability | The Malay toxicity data are served from a gated dataset named in the toxicity task’s configuration file (aisingapore/Safety-Toxicity-Detection) [2]; its public card lists no Malay split, although its file listing has a Malay folder, and on 30 September 2026 the data could not be downloaded without approved access | Independent reproduction now depends on being granted access |
| Confidence intervals | The ten highest of the 58 entries in the leaderboard page’s data share a 6.6-point interval range (Figure 4), and the five highest in its default view a 5.5-point range | Marginal intervals alone do not determine paired significance |
Under the official scorer, away from the zero floor, one corrected prediction moves the composite by 50/(3n) points on a safeguard subtask and by 50/n points on toxicity, where n is the size of the item’s class within its subtask (Figure 3). One of the 11 harmful items in the cultural-response subtask therefore moves the composite by 1.5 points, 15 times as much as one toxicity item (0.10) and about 20 times as much as one in-the-wild item (0.078). The cultural-prompt and cultural-response subtasks score only 71 distinct prompts (142 of the 1,572 scored items, 9.0 per cent), yet this is where single predictions move the composite most. These figures are for a single run, which is how we evaluate; published entries average eight runs per item, so an answer that changes in one of the eight runs moves an entry by an eighth of these amounts.
The data behind the leaderboard’s Malay page hold 58 open-weight entries of up to 200 billion parameters (dated 18 September 2026), of which the page’s default view displays 18; we use the full set. Its ten highest safety scores, from 46.22 down to 42.47, have 95% intervals that all share the range 40.8 to 47.4 points (Figure 4). Non-overlap of individual intervals is a conservative test for a difference, so we also approximated a 95% interval for each pairwise difference from the published half‑widths, treating entries as independent: every approximate interval includes zero, and the largest standardised difference is 0.93. This is an independence sensitivity calculation, not a paired test: the models share evaluation items, and the necessary covariance is unavailable in the published marginal intervals. The five highest entries in the default view likewise share a range (40.8 to 46.3); the same independence approximation includes zero for each pair. Further down that view some approximate intervals exclude zero, subject to the same covariance limitation. We therefore cannot determine from those marginal intervals which leading entries differ statistically. A paired test on shared items [5], with appropriate treatment of multiple comparisons, is needed.
Read the composite together with its uncertainty and subtask results. A small observed gap is inconclusive without an appropriate comparison; it is not automatically noise or evidence of equivalence. For models scored on shared items, report a paired interval on the difference. Read toxicity separately and validate the annotation policy against the intended use.
Practitioners should optimise behaviour on development data that reflects their own policy, use the benchmark as an external check, and report per-subtask results with intervals.
The paper’s own results are for one language, Malay; the single cross-language probe in Section 5.4 is a transfer check under its own serving configuration, not a multi-language safety evaluation, and no ordering between languages is claimed. What transfers is the structure of the problem. Item leverage is arithmetic: under a balanced-accuracy scorer, one changed prediction moves a subtask by an amount set by the size of the item’s class, so any language whose culturally grounded safety subtasks hold tens of items has the same exposure, and the expression in Section 5.1 can be applied to its subtask sizes directly. The benchmark family we assessed spans several Southeast Asian languages [1, 7], and comparable culturally grounded safety sets now exist for Indonesian languages [9].
Three reporting practices, items 1 to 3 of the checklist in Section 5.5, carry over unchanged: state which scorer and task set produced a score, declare the reasoning mode and request concurrency, and report an interval beside every score. What does not transfer is magnitude. How large each effect is depends on a language’s subtask sizes and annotation convention, and has to be measured language by language. For the same reason, our reading of the toxicity labels (Section 4.1) is a hypothesis to test in other languages, not a result about them.
The leverage arithmetic in Section 5.1 is a property of the scorer, not of Malay, so it should apply to any language whose culturally grounded safety subtasks are small. To test that expectation rather than assert it, we ran the same harness on a second regional language, Indonesian, which the model was not tuned for. No Indonesian-targeted text entered the fine-tuning mix, so the figures below are a cross-language transfer probe, not a within-language result, and no Indonesian-specific tuning or evaluation-driven development was performed.
Both checkpoints of one lineage were measured under a single shared
configuration: Oaica 35B-A3B Malay Safety v1.0 260923 and the
open-weight base Qwen3.6-35B-A3B that it is fine-tuned from. The harness
was SEA-HELM v1.3.0 on 20 of the 21 Indonesian tasks (the gated
syntax-criteria task could not be obtained for either
model), served as Q8_0 quantised weights under llama.cpp rather than the
vLLM bf16 stack of the Malay leaderboard, with sampling on and reasoning
off, over eight runs each. The serving stack and the run count therefore
differ from the Malay runs elsewhere in this paper, and these scores are
not comparable entry by entry with them.
Table 4: Indonesian SEA-HELM competency scores for two checkpoints of one lineage, under one shared configuration (SEA-HELM v1.3.0, 20 of 21 tasks, Q8_0 under llama.cpp, sampling on, reasoning off): mean and standard deviation (SD) over eight runs. The composite is the official safety score (Section 2); the overall figure is the equal-weight mean of the nine competencies above it. The two columns are not ordered (Section 5.2).
| Competency | Oaica 35B-A3B Malay Safety v1.0 260923 | Qwen3.6-35B-A3B |
|---|---|---|
| Multi-turn (LLM-judged) | 63.33 ± 2.45 | 84.95 ± 0.85 |
| Natural language understanding (NLU) | 77.25 ± 1.30 | 73.96 ± 4.16 |
| Safety (composite) | 57.16 ± 0.62 | 50.47 ± 9.46 |
| Natural language generation (NLG) | 55.10 ± 0.14 | 54.67 ± 0.19 |
| Natural language reasoning (NLR) | 87.45 ± 1.13 | 81.06 ± 5.87 |
| Linguistic diagnostics | 54.96 ± 6.58 | 44.05 ± 22.75 |
| Instruction following | 85.48 ± 1.63 | 87.62 ± 3.09 |
| Knowledge | 72.50 ± 5.07 | 67.92 ± 9.79 |
| Cultural | 76.80 ± 1.01 | 57.51 ± 14.95 |
| Overall (nine-competency mean) | 70.00 ± 1.59 | 66.91 ± 6.83 |
Two observations echo the Malay findings. First, the interval widths follow the item-leverage pattern of Section 5.1: the multi-turn and NLG subtasks are tightly determined (SD near one point), while linguistic diagnostics, knowledge and the cultural subtask carry SDs of five to about twenty-three points, the wide spread being what small item classes produce. Second, the composite again sits in the middle of its range, with the toxicity subtask low on both checkpoints, consistent with the hypothesis of Section 4.1 and with its cause left unresolved here as there.
Two of the eight base-model runs are anomalous: they collapse together across every classification subtask (toxicity accuracy 0.50 against about 0.71 elsewhere), which inflates that column’s SD. Over the six consistent runs the overall mean is 70.48 (SD 3.32), close to the fine-tuned checkpoint’s 70.00 (SD 1.59). The clearest and most reproducible separation is in the opposite direction, on the judge-rated multi-turn task, and there the two intervals do not overlap in this configuration. We report that gap rather than claim it as an advantage: the probe was not designed as a comparison, the two checkpoints are a fine-tune and its own base rather than independent entries, and no paired interval on the difference was computed (Section 5.1), so no ordering between the checkpoints is claimed. Both series are reported only to show that the measurement behaves in a second language as Section 5.1 predicts.
The findings above reduce to a short checklist for reporting any safety score, in any language. It describes what to publish beside a number, not how a model is built.
The checklist is implemented as a small open-source tool,
benchlint (github.com/sprapp-com/benchlint).
It lints a results report for the first four points above, flagging a
missing interval or sample size, ranking language with no test behind
it, false precision and single-run claims, and it computes item leverage
and interval estimates from item-level outcomes. It is language- and
model-agnostic, depends only on the Python standard library, and is
released under Apache-2.0.
This work sits where five lines of research meet: statistics for language-model evaluation, culturally grounded safety for Southeast Asian languages, annotator disagreement, the safety of reasoning models, and the validity of benchmarks and leaderboards.
Statistics for evaluations. Miller argues that evaluation results should carry error bars and be compared with paired analyses [3]. Bowyer et al. show that intervals based on the central limit theorem are unreliable on benchmarks with fewer than a few hundred items [4], and Wein et al. estimate significance for leaderboard differences from a single test set [5]. We apply these ideas to regional safety evaluation.
Regional and culturally grounded safety. SEA-HELM [1] evaluates Southeast Asian languages holistically; its leaderboard [2] has since added Malay, safeguard subtasks drawn from SEA-SafeguardBench [7], and bootstrap intervals over eight runs per model. Recent work builds culturally grounded safeguards and safety benchmarks for the region, including SEA-Guard [6], SEA-SafeguardBench [7], SEALGuard [8] and IndoSafety [9], and a joint testing exercise by several AI safety institutes covered Malay [10]. Our assessment complements this work by quantifying what the Malay safety track can resolve.
Annotator disagreement. Safety evaluation increasingly treats disagreement among annotators as information about genuine ambiguity rather than as noise [11], and recent modelling work separates such systematic disagreement from annotator error [12]. Automated safety annotations have also been compared with human ones, with imperfect agreement [13]. Our toxicity finding is consistent with this work.
Reasoning and safety. Studies of reasoning models examine the monitorability of reasoning traces and the safety of long chains of thought [14, 15]. We add an evaluation-side observation: switching reasoning on lowered the safety composite of both checkpoints (by 1.5 and 2.3 points, both historical paired intervals including zero), so the mode must be declared with every score.
Benchmark and leaderboard validity. Recent work asks what benchmarks measure and what their rankings support: Desai et al. examine validity across fifty-six AI benchmarks [16], and Yang and Chen study leaderboard claims under hidden model selection [17]. On SWE-bench Verified, Liu et al. report that paired tests do not distinguish any of the 29 adjacent pairs among the top thirty entries; on the larger Test split, some adjacent pairs are distinguishable [18]. Our assessment asks a related question of a regional-language safety track: which published differences exceed the benchmark’s own measurement uncertainty.
The evidence comes from one primary language with a single cross-language probe (Section 5.4), one benchmark family and, apart from the two reference points of Section 4.3, one base model, Qwen3.6-35B-A3B, under a compute budget that limited replication. Other model families appear only as published leaderboard entries in Sections 4.1 and 5. Several findings therefore may not generalise.
.oqm serving engine can pin the
model’s mixture-of-experts layers in host CPU RAM, filling the available
GPU memory automatically and offloading the remainder, so the same
weights run from laptop- and desktop-class GPUs at reduced context up to
datacenter servers, and higher-throughput engine configurations are
offered for API and cloud inference. The weights ship in our own
.oqm quantised container. These are deployment
characteristics, not validated performance guarantees, and they are
outside the measurements reported in this paper.Safety work handles harmful text and references to real communities, so the discipline constrains what is stored, published and released.
In this case study, some observed safety-score differences induced by evaluation settings were comparable to differences among closely related models. In the measurements reported here, aggregation moved one checkpoint by 5.3 points, request concurrency shifted the three cultural safeguard subtasks of one model by an amount worth about 1.8 composite points, and reasoning mode lowered the composite by 1.5 and 2.3 points, with paired intervals that include zero. By comparison, 12 training variants of one lineage spanned 3.3 points and differed pairwise by 1.2 points on average.
Toxicity performance remained low in the evaluated historical runs, but these results do not distinguish model limitations, annotation-policy mismatch, missing context or label error. The published intervals of leading entries overlap; without paired item-level outcomes we cannot determine which differences are statistically distinguishable. This case study supports reporting scorer definitions, inference settings, per-class errors and uncertainty. It does not establish comparative deployment safety, a universal benchmark ceiling or defects in the labels.
Early-stage Malay evaluation in this work used GlossoBench [19] (see Competing interests).
The author owns Oaica, which builds Malay language models and offers evaluation consulting. The study includes the company’s checkpoints, training variants, archived runs and serving measurements. Oaica also releases GlossoBench [19], an open evaluation harness used for early-stage Malay evaluation in this work. It describes SEA-HELM as its closest peer and is published under the GitHub account sprapp-com of Sprapp, the group of which Oaica is a part. The assessment of SEA-HELM in Sections 4 and 5 should be read with these interests in mind.
Oaica is a safety-focused AI company. It builds Malay language models and works with teams that need regional-language safety evaluation they can defend: audits of how a benchmark score was produced and how much of it the benchmark can support, reliability reviews of safety leaderboards, and on-premises Malay moderation consultation. De-identified evaluation tables, with their intervals, and item-leverage worksheets behind this paper are available on request. This paper: https://research.oaica.com/regional-language-llm-safety-malay-indonesian/. The Malay edition, which adds a section placing the company’s own checkpoints (its Section 6), is at https://research.oaica.com/malay-llm-safety-score/. Enquiries: [email protected] (general enquiries: [email protected]).