whisper-small-id

openai/whisper-small (244M params) fine-tuned for Indonesian: the most accurate CPU model, for RK3588 / Cortex-A76 boards and x86 CPUs. Code: github.com/porcupine-md/whisper-id. Lighter siblings for Cortex-A53: whisper-tiny-id, whisper-base-id.

Revisions: v2 (current, below), v1 (round 1: 147.5 h, FLEURS 11.4 / CV 10.3).

Training: 284.8 h (Common Voice 17 id, FLEURS id_id, YODAS2 id000 filtered with large-v3-turbo, YODAS labels = turbo transcripts), continued from v1 for 5000 steps; telephone / Opus / AAC / noise augmentation; 80% of batches with short encoder context (mel cut to the clip, as whisper.cpp -ac).

Results

WER / CER (%), full test sets, lowercase + punctuation stripped:

Test set whisper-small whisper-small-id v2 v2, short context
FLEURS id (684 utts) 19.4 / 5.8 9.9 / 3.0 9.9 / 3.0
Common Voice 17 id (3629 utts) 21.1 / 7.7 8.9 / 2.9 9.0 / 2.8

Telephone / Opus / AAC channel + short context (500 CV clips, training-time eval): 10.6 WER.

Latency, whisper.cpp q8_0, 4 s utterance:

CPU Settings ms / utterance RTF
RK3588, Cortex-A76 x4 (Orange Pi 5 Plus) taskset -c 4-7, -t 4 -ac 256 554 0.14
x86 AMD EPYC 7302P -t 4 -ac 256 598 0.15
x86 AMD EPYC 7302P -t 8 -ac 256 359 0.09
Cortex-A53 x4 (Orange Pi Zero3) -t 4 -ac 256 7200 1.8: not realtime, use tiny/base-id

Output is lowercase without punctuation.

Files

  • model.safetensors (fp16) + tokenizer/config: transformers
  • ggml-small-id-q8_0.bin (264 MB): whisper.cpp. On ARM, q8_0 beats q5_0 on speed.
  • ct2-int8/: CTranslate2 int8_float16 for faster-whisper on GPU (T4: 106 ms per 4 s utterance, single stream). On CPU prefer whisper.cpp: faster-whisper pads every clip to 30 s.

Usage

# -ac = encoder positions: 50 per second of audio, round up to a multiple of 64 (4 s -> 256, 8 s -> 448)
whisper-cli -m ggml-small-id-q8_0.bin -l id -t 4 -ac 256 -bs 1 -bo 1 -nt -f audio16k.wav
# GPU
from huggingface_hub import snapshot_download
from faster_whisper import WhisperModel
path = snapshot_download("maleo-ai/whisper-small-id", allow_patterns="ct2-int8/*")
model = WhisperModel(f"{path}/ct2-int8", device="cuda", compute_type="int8_float16")
segments, _ = model.transcribe("audio.wav", language="id", beam_size=1, vad_filter=True)
print(" ".join(s.text for s in segments))
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="maleo-ai/whisper-small-id")
print(asr("audio.wav", generate_kwargs={"language": "indonesian", "task": "transcribe"})["text"])

Limitations

  • Test sets are read speech; conversational audio will be harder.

Citation

@misc{irawan2026whispersmallid,
  author       = {Eka Tresna Irawan},
  title        = {whisper-small-id: Whisper small fine-tuned for Indonesian speech recognition},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/maleo-ai/whisper-small-id}}
}
Downloads last month
80
Safetensors
Model size
0.2B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for maleo-ai/whisper-small-id

Finetuned
(3795)
this model

Datasets used to train maleo-ai/whisper-small-id