Cleaned Polish Oscar corpus (128M lines)
Użyj poniższego opisu do cytowania zasobu albo wyeksportuj go w wybranym formacie:
Sopyła, Krzysztof, 2021,
Cleaned Polish Oscar corpus (128M lines), CLARIN-PL Repository,
http://hdl.handle.net/11321/845.
Autorzy
Strona projektu
Data
2021
Języki
Opis
Cleaned Polish Oscar corpus (part: 128M lines, 3.53 GB). Data was prepared with a few cleaning heuristics:
- remove sentences shorter than
- remove non-polish sentences
- remove ungrammatical sentences
- perform sentence tokenization and save each sentence in a new line, after each document the new line was added
Wydawca
Słowa kluczowe
Kolekcje