Language Descriptions
Stały URI dla kolekcjihttps://hdl.handle.net/11321/1016
Przeglądaj
Ostatnie zgłoszenia
Item type: Pozycja , Wordnet for Definition Augmentation with Encoder-Decoder Architecture(Global Wordnet Association, 2023-01-01) Wojtasik, Konrad; Janz, Arkadiusz; Alberski, Bartłomiej; Piasecki, MaciejData augmentation is a difficult task in Natural Language Processing. Simple methods that can be relatively easily applied in other domains like insertion, deletion or substitution, mostly result in changing the sentence meaning significantly and obtaining an incorrect example. Wordnets are potentially a perfect source of rich and high quality data that when integrated with the powerful capacity of generative models can help to solve this complex task. In this work, we use plWordNet, which is a wordnet of the Polish language, to explore the capability of encoder-decoder architectures in data augmentation of sense glosses. We discuss the limitations of generative methods and perform qualitative review of generated data samples.Item type: Pozycja , Wordnet-oriented Recognition of Derivational Relations(Global Wordnet Association, 2023-01-01) Walentynowicz, Wiktor; Piasecki, MaciejDerivational relations are an important element in defining meanings, as they help to explore word-formation schemes and predict senses of derivates (derived words). In this work, we analyse different methods of representing derivational forms obtained from WordNet – from quantitative vectors to contextual learned embedding methods – and compare ways of classifying the derivational relations occurring between them. Our research focuses on the explainability of the obtained representations and results. The data source for our research is plWordNet, which is the wordnet of the Polish language and includes a rich set of derivation examples.Item type: Pozycja , EnglishWordNet 2020: Improving and Extending aWordNet for English using an Open-Source Methodology(The European Language Resources Association (ELRA), 2020-05-01) McCrae, John; Rademaker, Alexandre; Rudnicka, Ewa; Bond, FrancisThe Princeton WordNet, while one of the most widely used resources for NLP, has not been updated for a long time, and as such a new project English WordNet has arisen to continue the development of the model under an open-source paradigm. In this paper, we detail the second release of this resource entitled “English WordNet 2020”. The work has focused firstly, on the introduction of new synsets and senses and developing guidelines for this and secondly, on the integration of contributions rom other projects. We present the changes in this edition, which total over 15,000 changes over the previous release.Item type: Pozycja , Towards a methodology for filtering out gaps and mismatches across wordnets: the case of noun synsets in plWordNet and Princeton WordNet(Global Wordnet Association, 2016-01-01) Rudnicka, Ewa; Witkowski, Wojciech; Grabowski, ŁukaszThis paper presents the results of large-scale noun synset mapping between plWordNet, the wordnet of Polish, and Princeton WordNet, the wordnet of English, which have shown high predominance of inter-lingual hyponymy relation over inter-synonymy relation. Two main sources of such effect are identified in the paper: differences in the methodologies of construction of plWN and PWN and cross-linguistic differences in lexicalization of concepts and grammatical categories between English and Polish. Next, we propose a typology of specific gaps and mismatches across wordnets and a rule-based system of filters developed specifically to scan all I(inter-lingual)-hyponymy links between plWN and PWN. The proposed system, it should be stressed, also enables one to pinpoint the frequencies of the identified gaps and mismatches.Item type: Pozycja , A (Non)-Perfect Match: Mapping plWordNet onto Princeton WordNet(Global Wordnet Association, 2021-01-01) Rudnicka, Ewa; Witkowski, Wojciech; Piasecki, MaciejThe paper reports on the methodology and final results of a large-scale synset mapping between plWordNet and Princeton WordNet. Dedicated manual and semi-automatic mapping procedures as well as interlingual relation types for nouns, verbs, adjectives and adverbs are described. The statistics of all types of interlingual relations are also provided.Item type: Pozycja , Lexical Perspective on Wordnet to Wordnet Mapping(Global Wordnet Association, 2018-01-01) Rudnicka, Ewa; Bond, Francis; Grabowski, Łukasz; Piasecki, Maciej; Piotrowski, TadeuszThe paper presents a feature-based model of equivalence targeted at (manual) sense linking between Princeton WordNet and plWordNet. The model incorporates insights from lexicographic and translation theories on bilingual equivalence and draws on the results of earlier synsetlevel mapping of nouns between Princeton WordNet and plWordNet. It takes into account all basic aspects of language such as form, meaning and function and supplements them with (parallel) corpus frequency and translatability. Three types of equivalence are distinguished, namely strong, regular and weak depending on the conformity with the proposed features. The presented solutions are language neutral and they can be easily applied to language pairs other than Polish and English. Sense-level mapping is a more finegrained mapping than the existing synset mappings and is thus of great potential to human and machine translation.Item type: Pozycja , Wordnet-based Evaluation of Large Distributional Models for Polish(Global Wordnet Association, 2018-01-01) Piasecki, Maciej; Czachor, Gabriela; Janz, Arkadiusz; Kaszewski, Dominik; Kędzia, PawełThe paper presents construction of large scale test datasets for word embeddings on the basis of a very large wordnet. They were next applied for evaluation of word embedding models and used to assess and compare the usefulness of different word embeddings extracted from a very large corpus of Polish. We analysed also and compared several publicly available models described in literature. In addition, several large word embeddings models built on the basis of a very large Polish corpus are presented.Item type: Pozycja , plWordNet 3.0 – Almost There(Global Wordnet Association, 2016-01-01) Piasecki, Maciej; Szpakowicz, Stan; Maziarz, Marek; Rudnicka, EwaIt took us nearly ten years to get from no wordnet for Polish to the largest wordnet ever built. We started small but quickly learned to dream big. Now we are about to release plWordNet 3.0-emo – complete with sentiment and emotions annotated – and a domestic version of PrincetonWordNet, larger thanWordNet 3.1 by nearly ten thousand newly added words. The paper retraces the road we travelled and talks a little about the future.Item type: Pozycja , Introduction to the special issue: On wordnets and relations(Springer, 2013-08-15) Piasecki, Maciej; Szpakowicz, Stan; Fellbaum, Christiane; Pedersen, Bolette SandfordThe present paper is concerned with the issues of wordnets and relations.Item type: Pozycja , Multisłownik: Linking plWordNet-based Lexical Data for Lexicography and Educational Purposes(Global Wordnet Association, 2018-01-01) Ogrodniczuk, Maciej; Bilińska, Joanna; Bronk, Zbigniew; Kieraś, WitoldMultisłownik is an automated integrator of Polish lexical data retrieved from multiple available online sources intended to be used in various scenarios requiring access to such data, most prominently dictionary creation, linguistic studies and education. In contrast to many available internet dictionaries Multisłownik is WordNet-centric, capturing the core definitions from Słowosieć synsets. The paper provides details of construction of the resource, discussed the difficulties related to linking different logical structures of underlying data and investigates two sample scenarios for using the resulting platform.Item type: Pozycja , WordnetLoom – a MultilingualWordnet Editing System Focused on Graph-based Presentation(Global Wordnet Association, 2018-01-01) Naskręt, Tomasz; Dziob, Agnieszka; Piasecki, Maciej; Saedi, Chakaveh; Branco, AntónioThe paper presents a new re-built and expanded, version 2.0 of WordnetLoom – an open wordnet editor. It facilitates work on a multilingual system of wordnets, is based on efficient software architecture of thin client, and offers more flexibility in enriching wordnet representation. This new version is built on the experience collected during the use of the previous one for more than 10 years of plWordNet development. We discuss its extensions motivated by the collected experience. A special focus is given to the development of a variant for the needs of MultiWordnet of Portuguese, which is based on a very different wordnet development model.Item type: Pozycja , A collaborative system for building and maintaining wordnets(Global Wordnet Association, 2019-07-01) Naskręt, TomaszA collaborative system for wordnet construction and maintenance is presented. Its key modules includeWordnetLoom editor, Wordnet Tracker and JavaScript Graph. They offer a number of functionalities that allow solving problems on every stage of building, editing and aligning wordnets by teams of lexicographers working in parallel. The experience collected in recent years has allowed us to refine applications and add new modules to provide the best user experience in a reliable and easily maintainable way.Item type: Pozycja , The chicken-and-egg problem in wordnet design: synonymy, synsets and constitutive relations(Springer (The Netherlands), 2013-04-18) Maziarz, Marek; Piasecki, Maciej; Szpakowicz, StanisławWordnets are built of synsets, not of words. A synset consists of words. Synonymy is a relation between words. Words go into a synset because they are synonyms. Later, a wordnet treats words as synonymous because they belong in the same synset. . . Such circularity, a well-known problem, poses a practical difficulty in wordnet construction, notably when it comes to maintaining consistency. We propose to make a wordnet a net of words or, to be more precise, lexical units. We discuss our assumptions and present their implementation in a steadily growing Polish wordnet. A small set of constitutive relations allows us to construct synsets automatically out of groups of lexical units with the same connectivity. Our analysis includes a thorough comparative overview of systems of relations in several influential wordnets. The additional synset-forming mechanisms include stylistic registers and verb aspect.Item type: Pozycja , Expanding WordNet with Gloss and Polysemy Links for Evocation Strength Recognition(Instytut Slawistyki Polskiej Akademii Nauk, 2020-12-01) Maziarz, Marek; Rudnicka, EwaEvocation — a phenomenon of sense associations going beyond standard (lexico)-semantic relations — is difficult to recognise for natural language processing systems. Machine learning models give predictions which are only moderately correlated with the evocation strength. It is believed that ordinary graph measures are not as good at this task as methods based on vector representations. The paper proposes a new method of enriching the WordNet structure with weighted polysemy and gloss links, and proves that Dijkstra’s algorithm performs equally as well as other more sophisticated measures when set together with such expanded structures.Item type: Pozycja , Towards Mapping Thesauri onto plWordNet(Global Wordnet Association, 2018-01-01) Maziarz, Marek; Piasecki, MaciejplWordNet, the wordnet of Polish, has become a very comprehensive description of the Polish lexical system. This paper presents a plan of its semi-automated integration with thesauri, terminological databases and ontologies, as a further necessary step in its development. This will improve linking of plWordNet into Linked Open Data, and facilitate applications in, e.g., WSD, keyword extraction or automated metadata generation. We present an overview of resources relevant to Polish and a plan for their linking to plWordNet.Item type: Pozycja , Testing agreement between lexicographers: A case of homonymy and polysemy(Global Wordnet Association, 2021-01-01) Maziarz, Marek; Bond, Francis; Rudnicka, EwaIn this paper we compare Oxford Lexico and Merriam Webster dictionaries with Princeton WordNet with respect to the description of semantic (dis)similarity between polysemous and homonymous senses that could be inferred from them. WordNet lacks any explicit description of polysemy or homonymy, but as a network of linked senses it may be used to compute semantic distances between word senses. To compare WordNet with the dictionaries, we transformed sample entry microstructures of the latter into graphs and crosslinked them with the equivalent senses of the former. We found that dictionaries are in high agreement with each other, if one considers polysemy and homonymy altogether, and in moderate concordance, if one focuses merely on polysemy descriptions. Measuring the shortest path lengths on WordNet gave results comparable to those on the dictionaries in predicting semantic dissimilarity between polysemous senses, but was less felicitous while recognising homonymy.Item type: Pozycja , Registers in the System of Semantic Relations in plWordNet(University of Tartu Press, 2014-01-01) Maziarz, Marek; Piasecki, Maciej; Rudnicka, Ewa; Szpakowicz, StanLexicalised concepts are represented in wordnets by word-sense pairs. The strength of markedness is one of the factors which influence word use. Stylistically unmarked words are largely contextneutral. Technical terms, obsolete words, “officialese”, slangs, obscenities and so on are all marked, often strongly, and that limits their use considerably. We discuss the position of register and markedness in wordnets with respect to semantic relations, and we list typical values of register. We illustrate the discussion with the system of registers in plWordNet, the largest Polish wordnet. We present a decision tree for the assignment of marking labels, and examine the consistency of the editing decisions based on that tree.Item type: Pozycja , plWordNet as the Cornerstone of a Toolkit of Lexico-semantic Resources(University of Tartu Press, 2014-01-01) Maziarz, Marek; Piasecki, Maciej; Rudnicka, Ewa; Szpakowicz, StanA wordnet is many things to many people: a graph of inter-related lexicalised concepts, a taxonomy, a thesaurus, and so on. A wordnet makes good sense as the mainstay of any deep automated semantic analysis of text. We have begun the construction of a multi-component, multi-use toolkit of natural language processing tools with plWordNet, a very large Polish wordnet, at its centre. The components will include plWordNet and its mapping onto an ontology (the upper level and elements of the middle level), a lexicon of proper names and a semantic valency lexicon. Some of those elements will be aligned with plWordNet, and there will be a mapping onto Princeton WordNet. Several challenging applications will show the utility of the toolkit in practice.Item type: Pozycja , Lexicalised and Non-lexicalized Multi-word Expressions inWordNet: a Cross-encoder Approach(Global Wordnet Association, 2023-01-01) Maziarz, Marek; Grabowski, Łukasz; Piotrowski, Tadeusz; Rudnicka, Ewa; Piasecki, MaciejFocusing on recognition of multi-word expressions (MWEs), we address the problem of recording MWEs in WordNet. In fact, not all MWEs recorded in that lexical database could with no doubt be considered as lexicalised (e.g. elements of wordnet taxonomy, quantifier phrases, certain collocations). In this paper, we use a cross-encoder approach to improve our earlier method of distinguishing between lexicalised and non-lexicalised MWEs found in WordNet using custom-designed rulebased and statistical approaches. We achieve F1-measure for the class of lexicalised word combinations close to 80%, easily beating two baselines (random and a majority class one). Language model also proves to be better than a feature-based logistic regression model.Item type: Pozycja , Adverbs in plWordNet: Theory and Implementation(Global Wordnet Association, 2016-01-01) Maziarz, Marek; Szpakowicz, Stan; Kaliński, MichałAdverbs are seldom well represented in wordnets. Princeton WordNet, for example, derives from adjectives practically all its adverbs and whatever involvement they have. GermaNet stays away from this part of speech. Adverbs in plWordNet will be emphatically present in all their semantic and syntactic distinctness. We briefly discuss the linguistic background of the lexical system of Polish adverbs. We describe an automated generator of accurate candidate adverbs, and introduce the lexicographic procedures which will ensure high consistency of wordnet editors’ decisions about adverbs.