Lexical and Conceptual Resources
Stały URI dla kolekcjihttps://hdl.handle.net/11321/1015
Przeglądaj
Ostatnie zgłoszenia
Item type: Pozycja , Register of multi-word expressions deleted from plWordNet after verification(Wrocław University of Science and Technology, 2024-08-29) Maziarz, Marek; Rudnicka, Ewa; Dziob, Agnieszka; Wieczorek, JustynaA dataset of multi-word expressions deleted from plWordNet after manual verification of their lexicality status.Item type: Pozycja , Polish multi-word lexical unit recognition(Wrocław University of Science and Technology, 2024-08-22) Rudnicka, Ewa; Maziarz, Marek; Grabowski, Łukasz; Pasternak, Simone; Przybysz, Zuzanna; Czerepowicka, MonikaA dataset of Polish multi-word expressions manually annotated with respect to their lexicality status. We show annotators' decisions with respect to two criteria: terminology (that is whether a given word combination can be classified as 'term', and 'paraphrase' (that is whether a given word combination can be can be easily paraphrased). In the last column, we present lexicographers' decision with respect to their lexicality status: "tak" - 'yes' means a given word combination is a multi-word lexical unit, "nie" - 'no' means it is not.Item type: Pozycja , plWordNet 4.2 (CLARIN-BIZ-START)(Wrocław University of Science and Technology, 2022-04-11) Dziob, Agnieszka; Tomasz, Naskręt; Maziarz, Marek; Wieczorek, Justyna; Dobrowolska-Pigoń, Marta; et al.plWordNet (Słowosieć) from Juli 2020, used as the main resources for word sense disambiguation tasks in 2020-2022; the database includes also the mapping to Priceton WordNet 3.1 and the PWN database.Item type: Pozycja , The subset of Polish pluralia tantum(Wroclaw University of Science and Technology, 2022-02-07) Rudnicka, Ewa; Grabowski, Łukasz; Paszke, Marta; Paszke, Kacper; Piasecki, MaciejThe subset of pluralia tantum extracted from the set of inter-lingual hyponyms between plWordNet and Princeton WordNet. The subset is tagged with the types of gaps and mismatches occurring between Polish and English and with the equivalence types wherever possible.Item type: Pozycja , Lexicalisation of Polish and English word combinations: two samples manually annotated (with collocation strength corpus statistics)(Wrocław University of Science and Technology, 2021-12-21) Dziob, Agnieszka; Grabowski, Łukasz; Kanclerz, Kamil; Kompa, Karolina; Maziarz, Marek; Piasecki, Maciej; Piotrowski, Tadeusz; Rudnicka, EwaWe analysed over 350 Polish and English word combinations (multi-word expressions, MWEs). Half of the sample was drawn from traditional dictionaries, while the other half was created by hand to represent free word combinations (i.e., MWEs not found in dictionaries, the information is given in the column "Status"). Syntactically these were noun phrases (NPs), either adjectives and nouns (A+N), or nouns and nouns (N+N), called 'bigrams'. We operationalised semantic compositionality by testing two custom-designed criteria, i.e., Intuition and Paraphrase, as well as by using statistical methods (selected measures of collocational strength, i.e. log-likelihood, PMI and Jaccard) for checking word order fixedness and word combination specificity. We also checked how long (in letters) the syntactic nucleus / its complement is (the measure highly correlated with word frequency, which is known as Zipf’s law (columns "AWL" and "HWL"). In the last column ("LCA") we give classification results obtained from Latent Class Analysis.Item type: Pozycja , plWordNet 4.5(Wrocław University of Science and Technology, 2021-07-23) Piasecki, MaciejPLWordNet ver. 4.5 is a lexico-semantic network that reflects the lexical system of the Polish language with projection to the English language. Słowosieć, Princeton Wordnet, EnWordnet together. It is now the largest wordnet in the world and is still growing.Item type: Pozycja , Mapping plWordNet 3.2 onto Linked Open Data - Manual Dataset(Wrocław University of Science and Technology, 2021-03-01) Biegalska, Amelia; Maziarz, Marek; Ohia-Nowak, MargaretThe mapping contains links to - AGROVOC - DDC - DIGIZAURUS - EUROVOC - GEMET - IATE - ICD10 - KABA - LCSH - MESH - RAMEAU - STERNIK - STW - UDC - UMLS - WIKIPEDIAItem type: Pozycja , Entry index and some auxiliary indexes to Linde's dictionary.(Formal Linguistics Department of Warsaw University, 2018) Bień, Janusz S.This is the archive of the mercurial repository formerly available at https://bitbucket.org/jsbien/ilindecsv. It contain the entry index and some auxiliary indexes to Linde's dictionary to be used with djview4poliqarp. For more information see Bień, Janusz S. “Elektroniczny indeks do słownika Lindego.” Kwartalnik Językoznawczy 2015, no. 3–4 (2018): 1–19. https://doi.org/10.14746/kj.2015.3-4.1. Bień, Janusz S. “Elektroniczne indeksy fiszek słownikowych.” Kwartalnik Językoznawczy 16, no. 2 (2018): 12. https://doi.org/10.14746/kj.2016.2.2. Bień, Janusz S., Janusz S. “Elektroniczne indeksy leksykograficzne i Djview4poliqarp.” Presented at the Seminarium „Przetwarzanie języka naturalnego”, Instytut Podstaw Informatyki PAN , Warszawa, January 10, 2018. https://www.slideshare.net/jsbien/jsb-i-linde181001ipi-117452985. Also https://www.youtube.com/watch?v=mOYzwpjTAf4.Item type: Pozycja , Indexes for djview4poliqarp(Formal Linguistics Department of Warsaw University, 2018) Bień, Janusz S.This is the archive of the mercurial repositories formerly available at https://bitbucket.org/jsbien/. They contain indexes to various resources in the DjVu format, in particular to the dictionary slips of some important Polish historical dictionaries. For more information see Bień, Janusz S. “Elektroniczne indeksy fiszek słownikowych.” Kwartalnik Językoznawczy 16, no. 2 (2018): 12. https://doi.org/10.14746/kj.2016.2.2. Bień, Janusz S., Janusz S. “Elektroniczne indeksy leksykograficzne i Djview4poliqarp.” Presented at the Seminarium „Przetwarzanie języka naturalnego”, Instytut Podstaw Informatyki PAN , Warszawa, January 10, 2018. https://www.slideshare.net/jsbien/jsb-i-linde181001ipi-117452985. Also https://www.youtube.com/watch?v=mOYzwpjTAf4.Item type: Pozycja , Walenty (2018-06-29)(Institute of Computer Science, Polish Academy of Sciences, 2018) Alberski, Bartłomiej; Andrejewicz, Jędrzej; Andrejewicz, Urszula; Andrzejczuk, Anna; Batko, Piotr; Brodzińska, Magdalena; Bukowiecka, Halina; Drabik, Lidia; Filipczak, Joanna; Grzeszak, Anna; Hajnicz, Elżbieta; Itoya, Bożena; Kaczmarska, Elżbieta; Kalużna-Gołąb, Marta; Kocyba, Natalia; Kozłowska, Matylda; Linsztet, Barbara; Łodzińska, Agnieszka; Maciejewska, Małgorzata; Norwa, Agnieszka; Opacki, Marcin; Patejuk, Agnieszka; Przepiórkowski, Adam; Rabiega-Wiśniewska, Joanna; Rosalska, Paulina; Skubida, Natalia; Skwarski, Filip; Stankiewicz, Anna; Sulich, Adrian; Szczyszek, Michał; Szymczak, Jakub; Świdziński, Marek; Wiśniakowska, Lidia; Woliński, Marcin; Wójcicka, Alicja; Zagajewska, Anna; Zawisławska, Magdalena; Zgondek, Maciej; Żochowska, Natalia; Żurowski, SebastianWalenty is a valence dictionary of Polish developed at the Institute of Computer Science, Polish Academy of Sciences (IPI PAN). The original formalism of Walenty was established by Filip Skwarski, Elżbieta Hajnicz, Agnieszka Patejuk, Adam Przepiórkowski, Marcin Woliński, Marek Świdziński, and Magdalena Zawisławska. It has been further developed by Elżbieta Hajnicz, Agnieszka Patejuk, Adam Przepiórkowski, and Marcin Woliński. The semantic layer has been developed by Elżbieta Hajnicz and Anna Andrzejczuk. The original seed of Walenty was provided by the automatic conversion, manually reviewed by Filip Skwarski, of the verbal valence dictionary used by the Świgra2 parser (6396 schemata for 1462 lemmata), which was in turn based on SDPV, the Syntactic Dictionary of Polish Verbs by Marek Świdziński (4148 schemata for 1064 lemmata). Afterwards, Walenty has been developed independently by adding new entries, syntactic schemata, in particular phraseological ones, and semantic frames. Walenty has been edited and compiled using the Slowal tool (http://zil.ipipan.waw.pl/Slowal) created by Bartłomiej Nitoń and Tomasz Bartosiak. The version of Walenty from 2018.06.29 contains 101 047 syntactic schemata and 28 321 semantic frames of 13022 verbs 4070 nouns, 950 adjectives and 200 nouns.Item type: Pozycja , Item type: Pozycja , Knowledge base of Polish conventionalized periphrastic nominal expressions(Institute of Computer Science, Polish Academy of Sciences, 2018) Nitoń, Bartłomiej; Ogrodniczuk, MaciejThe resource includes free Periphraser export with a knowledge base of Polish conventionalized periphrastic nominal expressions (i.e. phrases headed by a noun) together with their textually attested realizations. For instance, the database entry for the phrase ,,Robert Lewandowski'' in the referred resource will include the phrase ,,the Polish international'' while ,,pediatrics'' will be featured as ,,medical care for children''. Export is available in two formats provided by Periphraser - XML and CSV. Associated files include a free version of data (available on the CC BY-SA 4.0 License). In case you are interested in full (authenticated) export of Polish periphrastic data please contact resource contact person.Item type: Pozycja , enWordNet 1.0(Wrocław University of Technology, 2018-07-26) Piasecki, MaciejThe extension of Princeton WordNet built within the CLARIN-PL project. The attached file also contains the mapping to Open Multilingual Wordnet.Item type: Pozycja , plWordNet 4.0(Wrocław University of Technology, 2018-07-24) PIasecki, MaciejPLWordNet ver. 4.0 is a lexico-semantic network which reflects the lexical system of the Polish language with projection to the English language. Słowosieć, Princeton Wordnet, EnWordnet together the whole resource currently contains 506815 senses, 347564 synsets and over 1.5M relations and 361177 inter-lingual relations between lexical units. It is now the largest wordnet in the world and is still growing.Item type: Pozycja , Emotional Annotations Dictionary(ClarinPL, 2018-07-24) Zaśko-Zielińska, MonikaList of lexical units with emotional annotation extracted from Polish Wordnet (Słowosieć 4.0)Item type: Pozycja , Polish Dependency Bank(Institute of Computer Science, Polish Academy of Sciences, 2018-07-23) Wróblewska, AlinaPolish Dependency Bank (PDB) is the largest set of manually annotated dependency trees. PDB consists of more than 22K trees with 15.8 tokens per sentence on the average.Item type: Pozycja , Extended dictionary of named entities NELexicon connected with Linked Open Data(Wroclaw University of Science and Technology, 2018-07-19) Janz, Arkadiusz; Kocoń, JanThis resource contains Polish named entities connected with terminology from available resources within Linked Open Data (e.g. WordNet, DBPedia, Wikipedia, etc.).Item type: Pozycja , MWELexicon 1.1(Wrocław University of Technology, 2018-06-30) Dziob, Agnieszka; Kaliński, Michał; Maziarz, Marek; Piasecki, Maciej; Radziszewski, Adam; Szpakowicz, Stan; Wendelberger, MichałLexicon of 56,5k multi-word lexical units linked to plWordNet, together with description of their syntactic bahaviour obtained in constraint language (WCCL).Item type: Pozycja , Word combination lexicons(Wroclaw University of Science and Technology, 2017-07-11) Piasecki, Maciej10 lexicons comprising of several hundreds word combinations exhibiting various stages of lexicalisation (from free word combinations to fixed idioms) manually annotated by many linguists according to linguistic criteria (like semantic compositionality, syntactic idiosyncracy, terminological nature etc.) and intuitive sense of lexicalisation.Item type: Pozycja , Vector representations of polish words (Word2Vec method)(Wrocław University of Technology, 2016-11-07) Kędzia, Paweł; Czachor, Gabriela; Piasecki, Maciej; Kocoń, JanModel skip gram with vectors of length 100. Trained on kgr 10, a corpora with over 4 billion tokens. Data preprocessing involved segmentation, lemmatization and mophosyntactic disambiguation with MWE annotation.