Zaufany dostęp do polskich zasobów językowych
Dane językowe gotowe do badań
Odkrywaj, cytuj i deponuj polskie zasoby oraz narzędzia językowe. Zapewniamy im długoterminową ochronę i odpowiedzialne udostępnianie.
Odkrywaj
Przeglądaj kolekcję
Przeszukaj wszystkie zasobyKorpus Czterech Wieszczów
Wiersze wchodzące w skład Korpusu Czterech Wieszczów, zob. Tomasz Korpysz Anna Mędrzecka Ewa Mirkowska Marek Troszyński, „Korpus Czterech Wieszczów” – cyfrowy wymiar dziedzictwa narodowego. Założenia projektu, "Poradnik Językowy" 7/2022; T. Korpysz, A. Mędrzecka-Stefańska, O projekcie "Korpus Czterech Wieszczów", "Idiolekty" nr 2(2025).
Polish Drama Corpus
The Polish Drama Corpus (PolDraCor) contains 50 Polish-language plays from 1772–1939 encoded in TEI XML. The corpus was prepared by the Historical Pragmalinguistic Team at the University of Silesia in Katowice, with technical support from CLARIN-PL. This description was reconstructed from the project repository; the original DSpace description was unavailable.
Wordnet for Definition Augmentation with Encoder-Decoder Architecture
Data augmentation is a difficult task in Natural Language Processing. Simple methods that can be relatively easily applied in other domains like insertion, deletion or substitution, mostly result in changing the sentence meaning significantly and obtaining an incorrect example. Wordnets are potentially a perfect source of rich and high quality data that when integrated with the powerful capacity of generative models can help to solve this complex task. In this work, we use plWordNet, which is a wordnet of the Polish language, to explore the capability of encoder-decoder architectures in data augmentation of sense glosses. We discuss the limitations of generative methods and perform qualitative review of generated data samples.
Najczęściej oglądane
Najczęściej oglądane w ostatnim miesiącu
Korpus Czterech Wieszczów
Wiersze wchodzące w skład Korpusu Czterech Wieszczów, zob. Tomasz Korpysz Anna Mędrzecka Ewa Mirkowska Marek Troszyński, „Korpus Czterech Wieszczów” – cyfrowy wymiar dziedzictwa narodowego. Założenia projektu, "Poradnik Językowy" 7/2022; T. Korpysz, A. Mędrzecka-Stefańska, O projekcie "Korpus Czterech Wieszczów", "Idiolekty" nr 2(2025).
MultiCo-Hub: a corpus of multimodal enrichments with motion-trajectory annotation
MultiCo-Hub is a multimodal dataset including 11 zipped subsets (henceforth: sessions) of time-aligned audio, video and motion-capture–derived BVH data, together with multi-layered Annotation Pro files (ANTx) extended with automatically extracted motion-trajectory layers. The dataset includes a dedicated training session demonstrating body movement (TESM_001). The video and audio files included in remaining 10 sessions are derived from the MultiCo corpus (http://hdl.handle.net/11321/942). The original MultiCo sessions were enriched by means of: - full audio, video and BVH streams synchronization to enhance precise multimodal analysis; - motion-capture (BVH) data normalization, conversion, and integration directly into the annotation files as layers describing trajectories of selected body parts (positions, speeds, gesture-space coordinates). Furthermore, for each session, the corpus provides a composite multi-view video file showing all four camera angles simultaneously. This makes the dataset easier to inspect and substantially more accessible for users working on standard-performance computers. MultiCo-Hub offers a compact, ready-to-use resource for research and education in the areas of speech–gesture coordination, gesture space, temporal properties of movement, communicative alignment of interlocutors, and multimodal interaction. Export to common formats (TextGrid, EAF, CSV, etc.) is supported via Annotation Pro, facilitating downstream statistical analysis, visualization and interoperability. The MultiCo-Hub set also served as input for developing a set of R and C# applications and scripts that support the analysis and visualization of gesture space, temporal movement properties, and communicative alignment in dialogue.
Polish-Ukrainian Parallel Corpus
Polish-Ukrainian Parallel Corpus
Zaufanie i transparentność
Odpowiedzialny, długoterminowy dostęp
Nasze procedury obejmują ocenę depozytów, integralność danych, trwałą identyfikację, dostęp, ochronę i ciągłość działania. Repozytorium posiada certyfikat CoreTrustSeal na okres od 30 stycznia 2026 r. do 29 stycznia 2029 r.
Konsorcjum





