> For the complete documentation index, see [llms.txt](https://twbvoiceplaybook.clearglobal.org/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://twbvoiceplaybook.clearglobal.org/2.-what-is-voice-data/2.2-data-imbalance-and-low-resource-languages.md).

# 2.2 Data imbalance & low-resource languages

There is a serious imbalance in the [voice data](#user-content-fn-1)[^1] available for different languages. This is a big problem for the development of speech technology:

<table data-view="cards"><thead><tr><th></th><th></th></tr></thead><tbody><tr><td><strong>High-resource languages</strong></td><td>With languages like English, Mandarin, Spanish, and French, there are thousands of hours of voice data available. This results in very accurate speech technologies.</td></tr><tr><td><strong>Mid-resource languages</strong></td><td>With languages that have moderate populations of speakers or that are important in economic terms, there may be some commercial speech technology support. However, they still don’t have the amount of voice <a data-footnote-ref href="#user-content-fn-2">datasets</a> that are available for high-resource languages. Examples include Czech, Hindi, Polish, Tamil, Thai and Swahili.</td></tr><tr><td><strong>Low-resource languages</strong></td><td>For most of the world's 7,000+ languages, there is very little or no voice data available to develop technology. This is true even though millions of people speak these languages across the world. Kannada, Sindhi, and Tajik, for example, are “under-resourced” languages. Gondi, Khasi, and Santhali are “no-resource” languages as there is almost no data available for them.</td></tr></tbody></table>

The result of this imbalance is a "[digital language divide](#user-content-fn-3)[^3]". Speakers of low-resource languages cannot make use of voice technologies. Many of the languages spoken in areas affected by disaster or conflict are low-resource languages.

Some of the results of this imbalance:

* communities are not able to use digital services
* people cannot access key information during crises
* there is very little documentation of languages in danger of dying out
* global inequality is made worse

[^1]: **Voice data:** Audio recordings of human speech. These recordings capture the acoustic features of spoken language, such as pronunciation, speaking patterns, and rhythm.

[^2]: **Dataset:** A collection of information that has been organized for use. A **voice dataset** is a collection of voice recordings (paired with transcription) with additional information (metadata) such as gender, age of the person recording to give more information on how the data set is constructed and to avoid bias. It is for use in research and for training or improving voice models.

[^3]: **Digital language divide:** When it comes to access to technology, there is a big gap between people who speak high-resource languages and those who speak low-resource languages. It means that many people cannot enjoy the benefits of digitalization.
