diacnet-1.0 — Diacritic Restoration
A single byte-level (ByT5) model that restores diacritics for 10 languages: Yoruba, Vietnamese, Igbo, Hausa, Polish, Turkish, Portuguese, Spanish, French, Italian.
Model: olaverse/diacnet-1.0
· Fine-tuned from google/byt5-small · Apache-2.0
⚠️ Works best on single sentences or short passages (trained on sentence-length input, median 58 bytes). Longer text is automatically split on sentence boundaries below.
Target language
16 256
Examples
Known limitation: Yoruba CER (0.11) is notably higher than the other 9 languages (0.01–0.06), largely due to tone ambiguity in short inputs. See the model card for the full benchmark table.