Skip to main navigation Skip to search Skip to main content

Korean part-of-speech tagging based on morpheme generation

  • Hyun Je Song
  • , Seong Bae Park*
  • *Corresponding author for this work
    • Kyung Hee University

    Research output: Contribution to journalJournal articlepeer-review

    Abstract

    Two major problems of Korean part-of-speech (POS) tagging are that the word-spacing unit is not mapped one-to-one to a POS tag and that morphemes should be recovered during POS tagging. Therefore, this article proposes a novel two-step Korean POS tagger that solves the problems. This tagger first generates a sequence of lemmatized and recovered morphemes that can be mapped one-to-one to a POS tag using an encoder-decoder architecture derived from a POS-tagged corpus. Then, the POS tag of each morpheme in the generated sequence is finally determined by a standard sequence labeling method. Since the knowledge for segmenting and recovering morphemes is extracted automatically from a POS-tagged corpus by an encoderdecoder architecture, the POS tagger is constructed without a dictionary nor handcrafted linguistic rules. The experimental results on a standard dataset show that the proposed method outperforms existing POS taggers with its state-of-the-art performance.

    Original languageEnglish
    Article numberA41
    JournalACM Transactions on Asian and Low-Resource Language Information Processing
    Volume19
    Issue number3
    DOIs
    StatePublished - 2020.01.9

    Keywords

    • Morpheme generation
    • Morphologically complex languages
    • Part-of-speech tagging

    Quacquarelli Symonds(QS) Subject Topics

    • Computer Science & Information Systems

    Fingerprint

    Dive into the research topics of 'Korean part-of-speech tagging based on morpheme generation'. Together they form a unique fingerprint.

    Cite this