Użyj poniższego opisu do cytowania zasobu albo wyeksportuj go w wybranym formacie:
Dziob, Agnieszka; et al., 2021, Lexicalisation of Polish and English word combinations: two samples manually annotated (with collocation strength corpus statistics), CLARIN-PL Repository, http://hdl.handle.net/11321/853.
dc.contributor.authorDziob, Agnieszka
dc.contributor.authorGrabowski, Łukasz
dc.contributor.authorKanclerz, Kamil
dc.contributor.authorKompa, Karolina
dc.contributor.authorMaziarz, Marek
dc.contributor.authorPiasecki, Maciej
dc.contributor.authorPiotrowski, Tadeusz
dc.contributor.authorRudnicka, Ewa
dc.date.accessioned2021-12-21T13:44:42Z
dc.date.available2021-12-21T13:44:42Z
dc.date.issued2021-12-21
dc.descriptionWe analysed over 350 Polish and English word combinations (multi-word expressions, MWEs). Half of the sample was drawn from traditional dictionaries, while the other half was created by hand to represent free word combinations (i.e., MWEs not found in dictionaries, the information is given in the column "Status"). Syntactically these were noun phrases (NPs), either adjectives and nouns (A+N), or nouns and nouns (N+N), called 'bigrams'. We operationalised semantic compositionality by testing two custom-designed criteria, i.e., Intuition and Paraphrase, as well as by using statistical methods (selected measures of collocational strength, i.e. log-likelihood, PMI and Jaccard) for checking word order fixedness and word combination specificity. We also checked how long (in letters) the syntactic nucleus / its complement is (the measure highly correlated with word frequency, which is known as Zipf’s law (columns "AWL" and "HWL"). In the last column ("LCA") we give classification results obtained from Latent Class Analysis.
dc.identifier.urihttp://hdl.handle.net/11321/853
dc.language.isopol
dc.language.isoeng
dc.publisherWrocław University of Science and Technology
dc.rightsCreative Commons - Attribution 4.0 International (CC BY 4.0)
dc.rights.labelCC
dc.rights.urihttps://creativecommons.org/licenses/by/4.0/
dc.subjectmulti-word units
dc.subjectmulti-word expressions
dc.subjectMWE detection
dc.subjectlexical semantics
dc.subjectlexicography
dc.subjectsemantic compositionality
dc.titleLexicalisation of Polish and English word combinations: two samples manually annotated (with collocation strength corpus statistics)
dc.typelexicalConceptualResource
local.contact.personMarek Maziarz marek.maziarz@pwr.edu.pl Wrocław University of Science and Technology
local.files.count1
local.files.size8316
local.has.filesyes
local.language.namePolish
local.language.nameEnglish
local.size.info350 multiWordUnits
metashare.ResourceInfo#ContentInfo.detailedTypewordList
metashare.ResourceInfo#ContentInfo.mediaTypetext
Ten zasób jestCCi został udostępniony na licencji:Creative Commons - Attribution 4.0 International (CC BY 4.0)
Attribution Required

Pliki w tym zasobie

Nazwa
mwe.7z
Rozmiar
8.12 KB
Format
application/octet-stream
Opis
Unknown
Suma kontrolna MD5
23364d8a6fa6a166abff8927f4c13b8a
Preview
  Podgląd pliku
    The file preview has not been generated yet. Please try again later or contact the system administrator