Corpora

Stały URI dla kolekcjihttps://hdl.handle.net/11321/4

Przeglądaj

Ostatnie zgłoszenia

Teraz wyświetlane 1 - 20 z 269
  • Item type: Pozycja ,
    Korpus Czterech Wieszczów
    (Wrocław University of Science and Technology, 2026-05-07) Marek, Troszyński; Tomasz, Korpysz; Ewa, Mirkowska; Anna, Mędrzecka-Stefańska; Marcin, Oleksy; Tomasz, Bernaś; Tomasz, Naskręt; Maciej Piasecki
    Wiersze wchodzące w skład Korpusu Czterech Wieszczów, zob. Tomasz Korpysz Anna Mędrzecka Ewa Mirkowska Marek Troszyński, „Korpus Czterech Wieszczów” – cyfrowy wymiar dziedzictwa narodowego. Założenia projektu, "Poradnik Językowy" 7/2022; T. Korpysz, A. Mędrzecka-Stefańska, O projekcie "Korpus Czterech Wieszczów", "Idiolekty" nr 2(2025).
  • Item type: Pozycja ,
    Polish Drama Corpus
    (University of Silesia in Katowice, 2026-07-07) Pastuch, Magdalena; Mitrenga, Barbara; Wąsińska, Kinga
    The Polish Drama Corpus (PolDraCor) contains 50 Polish-language plays from 1772–1939 encoded in TEI XML. The corpus was prepared by the Historical Pragmalinguistic Team at the University of Silesia in Katowice, with technical support from CLARIN-PL. This description was reconstructed from the project repository; the original DSpace description was unavailable.
  • Item type: Pozycja ,
    DiPSS - longitudinal corpus of drift in Polish students of Spanish
    (Adam Mickiewicz University, Poznań, 2025-11-30) Sawicka-Stępińska, Brygida; Sypiańska, Jolanta
    The DiPSS corpus (part 1) is a longitudinal speech resource documenting the phonetic productions of L1 Polish students learning L2 English and L3 Spanish. It includes recordings from first year Spanish philology students across five testing points over two academic years, capturing word-initial stops (lenis and fortis), vowels (e, o, u, a), rhotics ({rr}) and approximants ([β, ð, ɣ]). The corpus integrates rich metadata including L2/L3 proficiency, language aptitude (LLAMA, Meara & Rogers, 2019), and age of onset for foreign languages, allowing for longitudinal and cross-linguistic analyses. DiPSS is designed as an open-access resource suitable for research in L1 drift, cross-linguistic influence, speech production and multilingual acquisition. Its detailed annotation, metadata and longitudinal structure result in a valuable tool for both linguistic research and computational modeling. The task consisted in reading words presented on auto-advancing slides in Polish, Spanish, and English. Instructions for the entire task were delivered in Polish. Prior to the Spanish and English sets of target words, participants received a written instruction along with a brief audio prompt in the respective language to establish the appropriate language mode. Audio was captured using the AKG C4000 microphone connected to a computer via a Focusrite Scarlett 2i2 audio interface and recorded using Audacity software, version 3.4.2. Data were collected from 28 speakers across testing times 1–4, and 22 speakers across testing times 1–5. The testing times correspond to: T1: October, year 1, during the opening week of the program, T2: November, year 1, after approximately five full weeks of instruction, T3: February, year 1, at the end of the first semester, T4: June, year 1, at the end of the first academic year, T5: June-September, year 2, at the end of the second year of studies. Metadata corresponding to the speakers include the following information: A: Sociodemographic data: speaker ID, gender, age B: Language background: self-reported L1, L2 and L3, level of Spanish: (A - absolute beginners, B - false beginners, C - advanced learners) C and D: L2 and L3 profile (self-reported proficiency, age of onset of formal education, age of exposure to naturalistic speech, stay in Spanish/English speaking countries for longer than a month, weekly exposure to naturalistic speech) E: Proficiency and language aptitude test results. The DiPSS corpus consists of five packages (T1-T5) of recordings with forced-aligned three-tier annotation in TextGrid, performed using WebMAUS Basic (Kisler, T. et al. 2017). Each package corresponds to one testing time and contains three sets of data: Polish, Spanish, and English. Packages T1-T4 each include 28 recordings per language, with corresponding TextGrid files. Package T5 includes 22 recordings per language, also with their corresponding TextGrid files. In total, the corpus comprises 402 pairs of WAV and TextGrid files from 28 speakers. The total recording time is approximately 20 hours, and the complete corpus size is 2.5 GB. The recordings in the released DiPSS corpus part 1 cover data collected in mid-2020s. The labels of the recordings adhere to a structured format: SPEAKER ID_TESTING TIME_LANGUAGE, wherein: SPEAKER ID corresponds to a unique speaker ID consisting of 6 characters, TESTING TIME corresponds to one of the five recording sessions (T1, T2, T3, T4, T5), LANGUAGE corresponds to the language in which the task was recorded (PL – Polish, ES – Spanish, EN – English). The data were processed using the server infrastructure developed within "Digital Research Infrastructure for the Arts and the Humanities" (POIR.04.02.00-00-D006/20).
  • Item type: Pozycja ,
    ROP
    (ROP, 2025-07-15) Xaw, Cul
    ROP
  • Item type: Pozycja ,
    Corpus of Nineteenth-Century French Texts on Palingenesis
    (Marta Sukiennicka, 2025-04-10) Sukiennicka, Marta
    A tagged corpus of nineteenth-century French texts related to the concept of palingenesis, constructed from the digital collections of the French National Library "Gallica". The corpus is annotated for discourse type (scientific, philosophical, literary), discipline (medicine, astronomy, politics, literary criticism, etc.), textual genre (treatise, essay, study, novel, poem, etc.), function (scientific concept, religious concept, metaphor, etc.), and tone of the concept's usage (serious, ironic, polemical, etc.). Each entry includes a concise interpretive summary of the term’s meaning within the given work, as well as a selection of representative or significant excerpts.
  • Item type: Pozycja ,
    MultiCo
    (Adam Mickiewicz University, Poznań, 2023-12-31) Karpiński, Maciej; Katarzyna, Klessa; Ewa, Jarmołowicz-Nowikow; Janusz, Taborek; Brygida, Sawicka-Stępińska; Michał, Piosik
    The MultiCo multimodal corpus is one of the outcomes of the project "Digital Research Infrastructure for the Humanities and Arts Studies DARIAH-PL." This project was funded by POIR 4.2 of the European Regional Development Fund from 2021 to 2023 and was carried out by a consortium of academic institutions across Poland with Adam Mickiewicz University, Poznan as a member of the consortium. The MultiCo multimodal corpus was developed at the Faculty of Modern Languages of Adam Mickiewicz University in Poznań. The motivation behind creating the corpus stems from contemporary research on interpersonal communication. The studies confirm that in order to understand and model the multifaceted process of communication, it's essential to study and describe not only speech but also other components of communication, such as gestures, facial expressions, and body posture. The MultiCo corpus was designed to support and facilitate this type of research approach. The corpus contains over 15 hours of recordings and consists of three sections: - Monologs representing persuasion in parliamentary speeches and motivational talks (TEDex), - Dialogs based on task-oriented activities recorded in a lab setting, - Multilogs illustrating discussions with multiple participants, exemplified by conversations on current sports events (TVP Sport 4-4-2). The monolog and multilog sections are based on materials available in public media or archives, while the dialog section includes task-oriented dialogs originally designed and recorded specifically for this resource.
  • Item type: Pozycja ,
    DN XXI 213 (trial corpus)
    (SWPS University, 2024-06-23) Jaworska, Julia; Jaworska, Julia; Jaworska, Julia
    This is a trial corpus.
  • Item type: Pozycja ,
    DN XXI 213 (trial corpus 2)
    (SWPS University, 2024-06-24) Jaworska, Julia
    This is a trial corpus.
  • Item type: Pozycja ,
    Korpus przemówień przedwyborczych Baracka Obamy
    (USWPS, 2024) Szwed, Marcin
    Korpus tekstowy przemówień Baracka Obamy z lat 2006-2015.
  • Item type: Pozycja ,
    The LnNor Corpus: A spoken multilingual corpus of non-native and native Norwegian, English and Polish (Part 2)
    (Adam Mickiewicz University, 2024-05-15) Wrembel, Magdalena; Hwaszcz, Krzysztof; Pludra, Agnieszka; Skałba, Anna; Weckwerth, Jarosław; Malarski, Kamil; Cal, Zuzanna; Kędzierska, Hanna; Czarnecki-Verner, Tristan; Balas, Anna; Kaźmierski, Kamil; Żychliński, Sylwiusz; Gruszecka, Justyna
    The LnNor corpus was created as part of the data collection in two projects: CLIMAD (Crosslinguistic influence in multilingualism across domains: phonology and syntax) and ADIM (Across-domain Investigations in Multilingualism: Modeling L3 Acquisition in Diverse Settings), led by Prof. Magdalena Wrembel at Adam Mickiewicz University in Poznań, Poland and by Prof. Marit Westergaard at the Arctic University of Norway, from December 2021 to April 2024 with funding from the National Science Centre (NCN) in Poland and Norway Grants. The CLIMAD and ADIM projects explored cross-linguistic influence (CLI) in the acquisition, processing, and use of a third language (L3/Ln) across various language domains and focused on different settings and stages of acquisition from a multilingual perspective. A range of sophisticated methodologies, such as perception and production tests, grammaticality judgement tasks and online brain imaging techniques like EEG, were leveraged to unravel the intricacies of multilingual processing. By capturing real-time insights into the interplay of cross-linguistic influences, the projects not only provided valuable contributions to the understanding of L3/Ln acquisition but also advanced theoretical frameworks in this field. Corpus data collection covered a broad range of speech elicitation tasks. The recordings consist of word, sentence and text reading, picture story description, video story retelling, spontaneous speech and socio-phonetic interviews in Polish, English and Norwegian. The corpus contains metadata based on the Language History Questionnaire (Li et al. 2020) such as age, gender, native languages, proficiency level, length of language exposure, age of onset. Data was collected from different groups of speakers: • L1 Polish learners of Norwegian as L3/Ln, attending Scandinavian studies at Poznań College of Modern Languages and the University of Szczecin (instructed learners); • L1 Polish learners of Norwegian as L3/Ln, living in Norway (naturalistic learners) • L1 English natives as controls • L1 Norwegian natives as controls Six types of speech tasks were recorded in Norwegian, English and Polish: • word reading • sentence reading • text reading (“The North Wind and the Sun”) • story telling (spontaneous) • picture description • picture story telling • video story telling • translation from Polish/English to Norwegian Metadata corresponding to the recordings include the following information: • speaker ID, age, gender, education, current residence, speaker status (instructed/naturalistic/native), native language, additional languages spoken • recording ID • language: PL (Polish), EN (English), NO (Norwegian) • status: L1, L2, L3/Ln • speech task: WR (word reading), SR1/2/... (sentence reading), TR1/2/... (text reading), PD (picture description), ST (story telling), VT (video story telling) • recording date, recording place, iteration, recording environment, recording device, type of microphone, noise level, etc. The labels of the recordings adhere to a structured format: PROJECT_SPEAKER ID_LANGUAGE STATUS_TASK, wherein: • PROJECT corresponds to the project within which the data were collected (A for ADIM, C for CLIMAD) • SPEAKER ID corresponds to a unique speaker ID consisting of 8 characters • LANGUAGE STATUS represents the language in which the task was recorded and its status for the speaker (e.g., L1PL, L2EN, L3NO) • TASK corresponds to the type of speech task recorded (e.g., TR, SR, WR, etc.) The LnNor corpus has been created to represent multilingual speech with a focus on L3/Ln Norwegian learners as well as native controls of Norwegian, English and Polish. The corpus is designed to study linguistic variation in learners acquiring Norwegian as a foreign language in instructed and naturalistic settings. Additionally, a subcorpus of native speech patterns is provided to serve as a benchmark, against which the learners' productions could be compared. Furthermore, part 2 of the corpus contains word alignment with orthographic transcriptions of speech to facilitate subsequent analyses across various linguistic domains. All speech samples were recorded with the use of Shure SM-35 unidirectional cardioid head-worn condenser microphones, using portable Marantz PMD620 solid state recorders with signal digitized at 48 kHz, 16-bit. This set-up was selected to minimize ambient noise and provide clear and focused recordings. The LnNOR corpus part 2 consists of 1671 annotated files from 164 speakers. The speakers included 113 L1 Polish, 33 L1 Norwegian and 18 L1 speakers of English. The total recording time is approximately 59 hours and the full size is 26 GB. The recordings in the released LnNor corpus part 2 cover data collected between 2023-2024.
  • Item type: Pozycja ,
    The LnNor Corpus: A spoken multilingual corpus of non-native and native Norwegian, English and Polish (Part 1)
    (Adam Mickiewicz University, 2024-01-31) Magdalena, Wrembel; Hwaszcz, Krzysztof; Agnieszka, Pludra; Skałba, Anna; Weckwerth, Jarosław; Walczak, Angelika; Sypiańska, Jolanta; Żychliński, Sylwiusz; Malarski, Kamil; Kędzierska, Hanna; Kaźmierski, Kamil; Gruszecka, Justyna; Dziubalska-Kolaczyk, Katarzyna; Czarnecki-Verner, Tristan; Cal, Zuzanna; Balas, Anna
    The LnNor corpus was created as part of the data collection in two projects: CLIMAD (Cross- linguistic influence in multilingualism across domains: phonology and syntax) and ADIM (Across-domain Investigations in Multilingualism: Modeling L3 Acquisition in Diverse Settings), led by Prof. Magdalena Wrembel at Adam Mickiewicz University in Poznań, Poland and by Prof. Marit Westergaard at the Arctic University of Norway, from December 2021 to April 2024 with funding from the National Science Centre (NCN) in Poland and Norway Grants. The CLIMAD and ADIM projects explored cross-linguistic influence (CLI) in the acquisition, processing, and use of a third language (L3/Ln) across various language domains and focused on different settings and stages of acquisition from a multilingual perspective. A range of sophisticated methodologies, such as perception and production tests, grammaticality judgement tasks and online brain imaging techniques like EEG, were leveraged to unravel the intricacies of multilingual processing. By capturing real-time insights into the interplay of cross-linguistic influences, the projects not only provided valuable contributions to the understanding of L3/Ln acquisition but also advanced theoretical frameworks in this field. Corpus data collection covered a broad range of speech elicitation tasks. The recordings consist of word, sentence and text reading, picture story description, video story retelling, spontaneous speech and socio-phonetic interviews in Polish, English and Norwegian. The corpus contains metadata based on the Language History Questionnaire (Li et al. 2020) such as age, gender, native languages, proficiency level, length of language exposure, age of onset. Data was collected from different groups of speakers: • L1 Polish learners of Norwegian as L3/Ln, attending Scandinavian studies at Poznań College of Modern Languages and the University of Szczecin (instructed learners); • L1 Polish learners of Norwegian as L3/Ln, living in Norway (naturalistic learners) • L1 English natives as controls • L1 Norwegian natives as controls • speakers of L2/L3/Ln English and L2/L3/Ln Norwegian with various L1 backgrounds Six types of speech tasks were recorded in Norwegian, English and Polish: • word reading • sentence reading • text reading (“The North Wind and the Sun”) • picture description • picture story telling • video story telling Metadata corresponding to the recordings include the following information: • speaker ID, age, gender, education, current residence, speaker status • (instructed/naturalistic/native), native language, additional languages spoken • recording ID • language: PL (Polish), EN (English), NO (Norwegian) • status: L1, L2, L3/Ln • speech task: WR (word reading), SR1/2/... (sentence reading), TR1/2/... (text reading), PD (picture description), ST (story telling), VT (video story telling) • recording date, recording place, iteration, recording environment, recording device, type of microphone, noise level, etc. The labels of the recordings adhere to a structured format: PROJECT_SPEAKER ID_LANGUAGE STATUS_TASK, wherein: • PROJECT corresponds to the project within which the data were collected (A for ADIM, C for CLIMAD) • SPEAKER ID corresponds to a unique speaker ID consisting of 8 characters • LANGUAGE STATUS represents the language in which the task was recorded and its status for the speaker (e.g., L1PL, L2EN, L3NO) • TASK corresponds to the type of speech task recorded (e.g., TR, SR, WR, etc.) The LnNor corpus has been created to represent multilingual speech with a focus on L3/Ln Norwegian learners as well as native controls of Norwegian, English and Polish. The corpus is designed to study linguistic variation in learners acquiring Norwegian as a foreign language in instructed and naturalistic settings. Additionally, a subcorpus of native speech patterns is provided to serve as a benchmark, against which the learners' productions could be compared. Furthermore, parts of the corpus contain word alignment with orthographic transcriptions of speech to facilitate subsequent analyses across various linguistic domains. All speech samples were recorded with the use of Shure SM-35 unidirectional cardioid head-worn condenser microphones, using portable Marantz PMD620 solid state recorders with signal digitized at 48 kHz, 16-bit. This set-up was selected to minimize ambient noise and provide clear and focused recordings. The LnNOR corpus part 1 consists of 1073 annotated files from 78 speakers. The speakers included 53 L1 Polish, 16 L1 Norwegian and 9 L1 speakers of other European languages. The total recording time is approximately 35 hours and the full size is 18 GB. The recordings in the released LnNor corpus part 1 cover data collected between 2021-2022.
  • Item type: Pozycja ,
    Mochnacki
    (CLARIN-PL, 2023-11-22) Mędrzecka, Anna; Bernaś, Tomasz
    korpus tekstów Mochnackiego
  • Item type: Pozycja ,
    Corpus of Russian Local Press of the Millennium Period (1996-2006)
    (Adam Mickiewicz University, 2023-04-14) Fedorushkov, Yury
    Corpus of Russian Local Press of the Millennium Period (1996-2006): selected archives (borders - from 1995/1996-2006) of two hundred and eighty (280) local newspapers from eighty-six (86) subjects of the Russian Federation (2005-2006): Oblasts (provinces), Republics, Krais (territories), Autonomous Okrugs (with a substantial ethnic minority), Federal cities, Autonomous Oblasts.
  • Item type: Pozycja ,
    PolEval 2019 Task 2: Lemmatization of proper names and multi-word phrases — train, tune and test data.
    (Wrocław University of Science and Technology, 2019-05-31) Marcińczuk, Michał
    The task consists in developing a tool for the lemmatization of proper names and multi-word phrases. The generated lemmas should follow the KPWr guidelines [https://clarin-pl.eu/dspace/handle/11321/625]. — poleval2019_task2_train.zip contains training data, — poleval2019_task2_tune.zip contains data used for the initial evaluation, — poleval2019_task2_test.zip contains data used for the final evaluation,
  • Item type: Pozycja ,
    Street name changes in Poznań, Słubice and Zbąszyń, Poland 1916-2018
    (Adam Mickiewicz University, 2022) Dobkiewicz, Patryk; Brzezińska, Anna Weronika; Fabiszak, Małgorzata
    The corpus presents a historical overview of street and place (park, bridge, square) name changes in the years 1916-2018 for three Polish cities: Poznań, Słubice and Zbąszyń. Included are the data for 2,582 streets in Poznań, 139 streets in Słubice and 105 streets in Zbąszyń, marked for the year of the introduction of a street name, the year when the name was changed or translated (if applicable), and the year when the name was removed (if applicable).
  • Item type: Pozycja ,
    StudEmo - corpus of consumer reviews annotated with emotions
    (Wrocław University of Science and Technology, 2022-05-20) Ngo, Anh; Candri, Argi; Ferdinan, Teddy; Kocoń, Jan; Korczyński, Wojciech
    Humans' emotional perception is subjective by nature, in which each individual could express different emotions regarding the same textual content. Existing datasets for emotion analysis commonly depend on a single ground truth per data sample, derived from majority voting or averaging the opinions of all annotators. We introduce a new non-aggregated dataset, namely StudEmo, that contains 5,182 customer reviews, each annotated by 25 people with intensities of eight emotions from Plutchik's model, extended with valence and arousal. We also propose three personalized models that use not only textual content but also the individual human perspective, providing the model with different approaches to learning human representations. The experiments were carried out as a multitask classification on two datasets: our StudEmo dataset and GoEmotions dataset, which contains 28 emotional categories. The proposed personalized methods significantly improve prediction results, especially for emotions that have low inter-annotator agreement.
  • Item type: Pozycja ,
    DiaBiz ASR benchmark
    (University of Lodz, 2022-04-27) Pęzik, Piotr; Adamczyk, Michał
    An evaluation report with accompanying datasets benchmarking the performance of commercially available ASR services of Polish on the DiaBiz corpus.
  • Item type: Pozycja ,
    Polish WSD Datasets
    (Wrocław University of Technology, 2022-04-11) Janz, Arkadiusz; Baran, Joanna; Oleksy, Marcin; Dziob, Agnieszka
    Data and code for the paper published at ICCS 2022: "A Unified Sense Inventory for Word Sense Disambiguation in Polish". The code is available at https://gitlab.clarin-pl.eu/team-semantics/wsd-research
  • Item type: Pozycja ,
    DiaBiz
    (University of Lodz, 2022) Pęzik, Piotr; Krawentek, Gosia; Karasińska, Sylwia; Wilk, Paweł; Rybińska, Paulina; Cichosz, Anna; Peljak-Łapińska, Angelika; Deckert, Mikołaj; Adamczyk, Michał
    DiaBiz corpus is a dialog corpus comprising recordings and annotated transcriptions of phone-based customer-agent interactions in several key business domains.
  • Item type: Pozycja ,
    DiaBiz.Kom sample 1.0
    (Wrocław University of Science and Technology, 2022-03-25) Oleksy, Marcin; Wieczorek, Jan; Domogała, Aleksandra; Drużyłowska, Dorota; Klyus, Julia; Wróż, Anita; Mikoś, Daria; Hwaszcz, Krzysztof; Kędzierska, Hanna
    DiaBiz.Kom sample is a sample of DiaBiz.Kom corpus, which is a dialog corpus comprising transcriptions of phone-based customer-agent interactions in several key business domains annotated with dialogue acts. Citation: Oleksy, M., Wieczorek, J., Drużyłowska, D., Klyus, J., Domogała, A., Hwaszcz, K., ... & Wróż, A. (2022, October). DiaBiz.Kom - towards a Polish Dialogue Act Corpus Based on ISO 24617-2 Standard. In Proceedings of the 29th International Conference on Computational Linguistics (pp. 3631-3638). (https://aclanthology.org/2022.coling-1.320) DiaBiz.Kom is based on DiaBiz corpus: http://docs.pelcra.pl/doku.php?id=diabiz