Użyj poniższego opisu do cytowania zasobu albo wyeksportuj go w wybranym formacie:
Wrembel, Magdalena; et al., 2024, The LnNor Corpus: A spoken multilingual corpus of non-native and native Norwegian, English and Polish (Part 2), CLARIN-PL Repository, http://hdl.handle.net/11321/932.
dc.contributor.authorWrembel, Magdalena
dc.contributor.authorHwaszcz, Krzysztof
dc.contributor.authorPludra, Agnieszka
dc.contributor.authorSkałba, Anna
dc.contributor.authorWeckwerth, Jarosław
dc.contributor.authorMalarski, Kamil
dc.contributor.authorCal, Zuzanna
dc.contributor.authorKędzierska, Hanna
dc.contributor.authorCzarnecki-Verner, Tristan
dc.contributor.authorBalas, Anna
dc.contributor.authorKaźmierski, Kamil
dc.contributor.authorŻychliński, Sylwiusz
dc.contributor.authorGruszecka, Justyna
dc.date.accessioned2024-05-17T12:27:24Z
dc.date.available2024-05-17T12:27:24Z
dc.date.issued2024-05-15
dc.descriptionThe LnNor corpus was created as part of the data collection in two projects: CLIMAD (Crosslinguistic influence in multilingualism across domains: phonology and syntax) and ADIM (Across-domain Investigations in Multilingualism: Modeling L3 Acquisition in Diverse Settings), led by Prof. Magdalena Wrembel at Adam Mickiewicz University in Poznań, Poland and by Prof. Marit Westergaard at the Arctic University of Norway, from December 2021 to April 2024 with funding from the National Science Centre (NCN) in Poland and Norway Grants. The CLIMAD and ADIM projects explored cross-linguistic influence (CLI) in the acquisition, processing, and use of a third language (L3/Ln) across various language domains and focused on different settings and stages of acquisition from a multilingual perspective. A range of sophisticated methodologies, such as perception and production tests, grammaticality judgement tasks and online brain imaging techniques like EEG, were leveraged to unravel the intricacies of multilingual processing. By capturing real-time insights into the interplay of cross-linguistic influences, the projects not only provided valuable contributions to the understanding of L3/Ln acquisition but also advanced theoretical frameworks in this field. Corpus data collection covered a broad range of speech elicitation tasks. The recordings consist of word, sentence and text reading, picture story description, video story retelling, spontaneous speech and socio-phonetic interviews in Polish, English and Norwegian. The corpus contains metadata based on the Language History Questionnaire (Li et al. 2020) such as age, gender, native languages, proficiency level, length of language exposure, age of onset. Data was collected from different groups of speakers: • L1 Polish learners of Norwegian as L3/Ln, attending Scandinavian studies at Poznań College of Modern Languages and the University of Szczecin (instructed learners) • L1 Polish learners of Norwegian as L3/Ln, living in Norway (naturalistic learners) • L1 English natives as controls • L1 Norwegian natives as controls Six types of speech tasks were recorded in Norwegian, English and Polish: • word reading • sentence reading • text reading (“The North Wind and the Sun”) • story telling (spontaneous) • picture description • picture story telling • video story telling • translation from Polish/English to Norwegian Metadata corresponding to the recordings include the following information: • speaker ID, age, gender, education, current residence, speaker status (instructed/naturalistic/native), native language, additional languages spoken • recording ID • language: PL (Polish), EN (English), NO (Norwegian) • status: L1, L2, L3/Ln • speech task: WR (word reading), SR1/2/... (sentence reading), TR1/2/... (text reading), PD (picture description), ST (story telling), VT (video story telling) • recording date, recording place, iteration, recording environment, recording device, type of microphone, noise level, etc. The labels of the recordings adhere to a structured format: PROJECT_SPEAKER ID_LANGUAGE STATUS_TASK, wherein: • PROJECT corresponds to the project within which the data were collected (A for ADIM, C for CLIMAD) • SPEAKER ID corresponds to a unique speaker ID consisting of 8 characters • LANGUAGE STATUS represents the language in which the task was recorded and its status for the speaker (e.g., L1PL, L2EN, L3NO) • TASK corresponds to the type of speech task recorded (e.g., TR, SR, WR, etc.) The LnNor corpus has been created to represent multilingual speech with a focus on L3/Ln Norwegian learners as well as native controls of Norwegian, English and Polish. The corpus is designed to study linguistic variation in learners acquiring Norwegian as a foreign language in instructed and naturalistic settings. Additionally, a subcorpus of native speech patterns is provided to serve as a benchmark, against which the learners' productions could be compared. Furthermore, part 2 of the corpus contains word alignment with orthographic transcriptions of speech to facilitate subsequent analyses across various linguistic domains. All speech samples were recorded with the use of Shure SM-35 unidirectional cardioid head-worn condenser microphones, using portable Marantz PMD620 solid state recorders with signal digitized at 48 kHz, 16-bit. This set-up was selected to minimize ambient noise and provide clear and focused recordings. The LnNOR corpus part 2 consists of 1671 annotated files from 164 speakers. The speakers included 113 L1 Polish, 33 L1 Norwegian and 18 L1 speakers of English. The total recording time is approximately 59 hours and the full size is 26 GB. The recordings in the released LnNor corpus part 2 cover data collected between 2023-2024.
dc.identifier.urihttp://hdl.handle.net/11321/932
dc.language.isonor
dc.language.isoeng
dc.language.isopol
dc.publisherAdam Mickiewicz University
dc.rightsCreative Commons - Attribution 4.0 International (CC BY 4.0)
dc.rights.labelCC
dc.rights.urihttps://creativecommons.org/licenses/by/4.0/
dc.source.urihttps://adim.web.amu.edu.pl/en/
dc.subjectL2 English
dc.subjectL3 Norwegian
dc.subjectL1 Polish
dc.subjectspoken data
dc.titleThe LnNor Corpus: A spoken multilingual corpus of non-native and native Norwegian, English and Polish (Part 2)
dc.typecorpus
local.contact.personKrzysztof Hwaszcz krzysztof.hwaszcz@uwr.edu.pl University of Wroclaw
local.demo.urihttps://adim.web.amu.edu.pl/en/lnnor-corpus/
local.files.count4
local.files.size845727108
local.has.filesyes
local.language.nameNorwegian
local.language.nameEnglish
local.language.namePolish
local.size.info26 gb
local.size.info59 hours
local.sponsor UMO-2019/34/H/HS2/00495 NCN GRIEG-1 project financed by EEA and Norway Grants 1. Across-domain Investigations in Multilingualism: Modeling L3 Acquisition in Diverse Settings (ADIM)
local.sponsor UMO-2020/37/B/HS2/00617  OPUS-19-HS financed by Polish National Science Centre 2. Cross-linguistic influence in multilingualism across domains: Phonology and syntax (CLIMAD)
metashare.ResourceInfo#ContentInfo.mediaTypeaudio

Kolekcje

Ten zasób jestCCi został udostępniony na licencji:Creative Commons - Attribution 4.0 International (CC BY 4.0)
Attribution Required

Pliki w tym zasobie

Nazwa
Description_LnNOR Corpus_part 2[68].pdf
Rozmiar
78.97 KB
Format
application/pdf
Opis
Suma kontrolna MD5
676491997084f08529583ea7c8027d45
Preview
  Podgląd pliku
    The file preview has not been generated yet. Please try again later or contact the system administrator
Nazwa
Metadata_Part 2.xlsx
Rozmiar
28.04 KB
Format
application/vnd.openxmlformats-officedocument.spreadsheetml.sheet
Opis
Suma kontrolna MD5
550b6b21e6d721354e28278053a2f68c
Preview
  Podgląd pliku
    The file preview has not been generated yet. Please try again later or contact the system administrator
Nazwa
textgrid+txt.zip
Rozmiar
20.31 MB
Format
application/zip
Opis
Suma kontrolna MD5
c17c29573f57a91d54e1fdb6749e2f4d
Preview
  Podgląd pliku
    The file preview has not been generated yet. Please try again later or contact the system administrator
Nazwa
FINAL_Corpus_Part 2 MP3.zip
Rozmiar
786.13 MB
Format
application/zip
Opis
Suma kontrolna MD5
39e4dfd0ee1f0cc072ee2ba06b06b3c6
Preview
  Podgląd pliku
    The file preview has not been generated yet. Please try again later or contact the system administrator