Cleaned Polish Oscar corpus (128M lines)

Użyj poniższego opisu do cytowania zasobu albo wyeksportuj go w wybranym formacie:
Sopyła, Krzysztof, 2021, Cleaned Polish Oscar corpus (128M lines), CLARIN-PL Repository, http://hdl.handle.net/11321/845.
Data
2021
Języki
Opis
Cleaned Polish Oscar corpus (part: 128M lines, 3.53 GB). Data was prepared with a few cleaning heuristics: - remove sentences shorter than - remove non-polish sentences - remove ungrammatical sentences - perform sentence tokenization and save each sentence in a new line, after each document the new line was added
Wydawca
Słowa kluczowe
Kolekcje