Użyj poniższego opisu do cytowania zasobu albo wyeksportuj go w wybranym formacie:
Kocoń, Jan, 2017, CorpoGrabber, CLARIN-PL Repository, http://hdl.handle.net/11321/403.
dc.contributor.authorKocoń, Jan
dc.date.accessioned2017-06-28T09:14:07Z
dc.date.available2017-06-28T09:14:07Z
dc.date.issued2017-06-28
dc.descriptionCorpoGrabber: The Toolchain to Automatic Acquiring and Extraction of the Website Content Jan Kocoń, Wroclaw University of Technology CorpoGrabber is a pipeline of tools to get the most relevant content of the website, including all subsites (up to the user-defined depth). The proposed toolchain can be used to build a big Web corpora of text documents. It requires only the list of the root websites as the input. Tools composing CorpoGrabber are adapted to Polish, but most subtasks are language independent. The whole process can be run in parallel on a single machine and includes the following tasks: downloading of the HTML subpages of each input page URL [1], extracting of plain text from each subpage by removing boilerplate content (such as navigation links, headers, footers, advertisements from HTML pages) [2], deduplication of plain text [2], removing of bad quality documents utilizing Morphological Analysis Converter and Aggregator (MACA) [3], tagging of documents using Wrocław CRF Tagger (WCRFT) [4]. Last two steps are available only for Polish. The result is a corpora as a set of tagged documents for each website. References [1] https://www.httrack.com/html/faq.html [2] J. Pomikalek. 2011. Removing Boilerplate and Duplicate Content from Web Corpora. Ph.D. Thesis. Masaryk University, Faculcy of Informatics. Brno. [3] A. Radziszewski, T. Sniatowski. 2011. Maca – a configurable tool to integrate Polish morphological data. Proceedings of the Second International Workshop on Free/Open-Source Rule-Based Machine Translation. Barcelona, Spain. [4] A. Radziszewski. 2013. A tiered CRF tagger for Polish. Intelligent Tools for Building a Scientific Information Platform: Advanced Architectures and Solutions. Springer Verlag.
dc.identifier.urihttp://hdl.handle.net/11321/403
dc.language.isopol
dc.language.isoeng
dc.publisherJan Kocoń
dc.rightsGNU LGPL 3.0
dc.rights.labelPUB
dc.rights.urihttp://www.gnu.org/licenses/lgpl.html
dc.subjectCorpoGrabber
dc.subjectcorpus
dc.subjectacquiring
dc.subjectweb scraping
dc.subjectcorpora builder
dc.titleCorpoGrabber
dc.typetoolService
local.contact.personJan Kocoń jan.kocon@pwr.edu.pl Wroclaw University of Science and Technology
local.files.count1
local.files.size5123
local.has.filesyes
local.language.namePolish
local.language.nameEnglish
local.sponsornationalFunds 6358/IA/119/2013 Ministry of Science and Higher Education (Poland) CLARIN-PL
metashare.ResourceInfo#ContentInfo.detailedTypetool
metashare.ResourceInfo#ResourceComponentType#ToolServiceInfo.languageDependenttrue
Ten zasób jestPublicznie dostępnyi został udostępniony na licencji:GNU LGPL 3.0

Dostęp do plików

Pliki w tym zasobie

Pobierz kompletny pakiet albo skopiuj gotowe polecenie do automatycznego pobierania.

Nazwa
corpograbber.zip
Rozmiar
5 KB
Format
application/zip
Opis
Unknown
Suma kontrolna MD5
c417bb22d82c3cc9aeb92c77687190a0
Preview
Podgląd pliku
  • Podgląd niedostępny

    Podgląd tego pliku nie został jeszcze wygenerowany. Spróbuj ponownie później albo skontaktuj się z administratorem systemu.