diacnet-1.0 — Diacritic Restoration

A single byte-level (ByT5) model that restores diacritics for 10 languages: Yoruba, Vietnamese, Igbo, Hausa, Polish, Turkish, Portuguese, Spanish, French, Italian.

Model: olaverse/diacnet-1.0 · Fine-tuned from google/byt5-small · Apache-2.0

⚠️ Works best on single sentences or short passages (trained on sentence-length input, median 58 bytes). Longer text is automatically split on sentence boundaries below.

Target language
16 256
Examples

Known limitation: Yoruba CER (0.11) is notably higher than the other 9 languages (0.01–0.06), largely due to tone ambiguity in short inputs. See the model card for the full benchmark table.