Użyj poniższego opisu do cytowania zasobu albo wyeksportuj go w wybranym formacie:
Sopyła, Krzysztof, 2021, Cleaned Polish Oscar corpus (128M lines), CLARIN-PL Repository, http://hdl.handle.net/11321/845.
dc.contributor.authorSopyła, Krzysztof
dc.date.accessioned2021-07-30T11:23:09Z
dc.date.available2021-07-30T11:23:09Z
dc.date.issued2021
dc.descriptionCleaned Polish Oscar corpus (part: 128M lines, 3.53 GB). Data was prepared with a few cleaning heuristics: - remove sentences shorter than - remove non-polish sentences - remove ungrammatical sentences - perform sentence tokenization and save each sentence in a new line, after each document the new line was added
dc.identifier.urihttp://hdl.handle.net/11321/845
dc.language.isopol
dc.publisherErmlab
dc.source.urihttps://github.com/Ermlab/PoLitBert/
dc.subjectcorpus
dc.titleCleaned Polish Oscar corpus (128M lines)
dc.typecorpus
local.contact.personKrzysztof Sopyła office@ermlab.com Ermlab
local.demo.urihttps://minio.clarin-pl.eu/ermlab/public/PoLitBert/corpus-oscar/corpus_oscar_2020-04-10_128M_lines.zip
local.files.count0
local.files.size0
local.has.filesno
local.language.namePolish
metashare.ResourceInfo#ContentInfo.mediaTypetext

Kolekcje