<journal article>
CEFR-based Lexical Simplification Dataset

Creator
Language
Publisher
Date
Source Title
Vol
First Page
Last Page
Conference
Publication Type
Access Rights
Rights
Related DOI
Related DOI
Related URI
Related HDL
Abstract This study creates a language dataset for lexical simplification based on Common European Framework of References for Languages (CEFR) levels (CEFR-LS). Lexical simplification has continued to be one ...of the important tasks for language learning and education.There are several language resources for lexical simplification that are available for generating rules and creating simplifiers using machine learning. However, these resources are not tailored to language education with word levels and lists of candidates tending to be subjective. Different from these, the present study constructs a CEFR-LS whose target and candidate words are assigned CEFR levels using CEFR-J wordlists and English Vocabulary Profile, and candidates are selected using an online thesaurus. Since CEFR is widely used around the world, using CEFR levels makes it possible to apply a simplification method based on our dataset to language education directly. CEFR-LS currently includes 406 targets and 4912 candidates. To evaluate the validity of CEFR-LS for machine learning, two basic models are employed for selecting candidates and the results are presented as a reference for future users of the dataset.show more

Hide fulltext details.

pdf LREC2018_238 pdf 195 KB 440  

Details

Record ID
Related URI
Related ISBN
Subject Terms
Created Date 2018.06.14
Modified Date 2018.06.14

People who viewed this item also viewed