Natural Language Processing

Course Objectives :

  1. Analysing the concept and techniques of Natural Language Processing based on Morphology and CORPUS.
  2. Mathematical foundations, Probability theory with Linguistic essentials such as syntactic and semantic analysis of text.
  3. The application of statistical learning methods and cutting-edge research models from deep learning.
  4. Applying the natural language and suitable modelling technique based on the structure.

Course Outcomes (CO)

  1. CO 1 - Understand the knowledge of complex language behaviour in terms of phonetics, morphology etc.
  2. CO 2 - Understand the semantic and pragmatics for text processing to compile and analyse the texts based on digestive approach.
  3. CO 3 - Apply Part-of-speech (POS) tagging for a given natural language and suitable modelling technique.
  4. CO 4 - Apply state of the art algorithms and techniques for text based processing of natural language with respect to morphology.

UNIT-I
Introduction to NLP, The classical tool kit, Knowledge in speech and Language processing, ambiguity and models and algorithm, Language and understanding, brief history, Regular Expressions, patterns, words, Text normalization, Minimum edit distance, Regular Language and FSAs, Raw Text Extraction and Tokenization, Extracting Terms from Tokens, Normalization.

UNIT-II
N-grams, Evaluating language model, Generalization and zeros, smoothing, kneser-Ney smoothing, huge language models and stupid back off. Perplexity’s relation to entropy, Inflection, Derivational Morphology, Finite-State Morphological Parsing, Lexical and Morphotactics, Morphological Parsing with Finite State Transducers, Combining FST Lexicon and rules.

UNIT-III
Methodological Preliminaries, Supervised Disambiguation: Bayesian Classification, an information theoretic approach, Dictionary-Based Disambiguation: Disambiguation based on sense, Thesaurus based disambiguation, Disambiguation based on translations in a second-language corpus.

UNIT - IV
Markov Model: Hidden Markov model Fundamentals, Probabilities of properties, Parameter estimation, Variants, Multiple input observation. The Information sources in Tagging: Markov model and taggers, Viterbi algorithm, Applying HMMs to POS tagging, Applications of Tagging.

Textbook(s):
  1. Daniel Jurafsky and James H. Martin, “Speech and Language Processing”, Prentice Hall.
  2. Christopher D. Manning and Hinrich Schutze, “Foundation of Natural Language Processing”, The MIT Press Cambridge.
References:
  1. James Allen, “Natural Language Understanding:, Pearson Publication.
  2. Nitin Indurkhya, Fred J. Damerau, “Handbook of Natural Language Processing”, CRC Press, 2010.

No comments:

Post a Comment