Cleaned Polish Oscar corpus (96M lines)

Użyj poniższego opisu do cytowania zasobu albo wyeksportuj go w wybranym formacie:
Sopyła, Krzysztof, 2021, Cleaned Polish Oscar corpus (96M lines), CLARIN-PL Repository, http://hdl.handle.net/11321/844.
Data
2021
Języki
Opis
Cleaned Polish Oscar corpus (part: 96M lines, 3.49 GB). Data was prepared with a few cleaning heuristics: - remove sentences shorter than - remove non-polish sentences - remove ungrammatical sentences - perform sentence tokenization and save each sentence in a new line, after each document the new line was added
Wydawca
Słowa kluczowe
Kolekcje