Statistical parsing of varieties of clinical Finnish

Veronika Laippala; Timo Viljanen; Antti Airola; Jenna Kanerva; Sanna Salanterä; Tapio Salakoski; Filip Ginter

doi:10.1016/j.artmed.2014.02.002

Wide-Coverage Efficient Statistical Parsing with CCG and Log-Linear Models

Computational Linguistics ◽

10.1162/coli.2007.33.4.493 ◽

2007 ◽

Vol 33 (4) ◽

pp. 493-552 ◽

Cited By ~ 125

Author(s):

Stephen Clark ◽

James R. Curran

Keyword(s):

Linear Models ◽

Parallel Implementation ◽

Discriminative Training ◽

Training Data ◽

Highly Efficient ◽

Independent Events ◽

Cluster Dynamic ◽

Statistical Parsing ◽

Parsing Algorithm ◽

Log Linear

This article describes a number of log-linear parsing models for an automatically extracted lexicalized grammar. The models are “full” parsing models in the sense that probabilities are defined for complete parses, rather than for independent events derived by decomposing the parse tree. Discriminative training is used to estimate the models, which requires incorrect parses for each sentence in the training data as well as the correct parse. The lexicalized grammar formalism used is Combinatory Categorial Grammar (CCG), and the grammar is automatically extracted from CCGbank, a CCG version of the Penn Treebank. The combination of discriminative training and an automatically extracted grammar leads to a significant memory requirement (up to 25 GB), which is satisfied using a parallel implementation of the BFGS optimization algorithm running on a Beowulf cluster. Dynamic programming over a packed chart, in combination with the parallel implementation, allows us to solve one of the largest-scale estimation problems in the statistical parsing literature in under three hours. A key component of the parsing system, for both training and testing, is a Maximum Entropy supertagger which assigns CCG lexical categories to words in a sentence. The supertagger makes the discriminative training feasible, and also leads to a highly efficient parser. Surprisingly, given CCG's “spurious ambiguity,” the parsing speeds are significantly higher than those reported for comparable parsers in the literature. We also extend the existing parsing techniques for CCG by developing a new model and efficient parsing algorithm which exploits all derivations, including CCG's nonstandard derivations. This model and parsing algorithm, when combined with normal-form constraints, give state-of-the-art accuracy for the recovery of predicate-argument dependencies from CCGbank. The parser is also evaluated on DepBank and compared against the RASP parser, outperforming RASP overall and on the majority of relation types. The evaluation on DepBank raises a number of issues regarding parser evaluation. This article provides a comprehensive blueprint for building a wide-coverage CCG parser. We demonstrate that both accurate and highly efficient parsing is possible with CCG.

Download Full-text

Fast statistical parsing of noun phrases for document indexing

10.3115/974557.974603 ◽

1997 ◽

Cited By ~ 20

Author(s):

Chengxiang Zhai

Keyword(s):

Noun Phrases ◽

Document Indexing ◽

Statistical Parsing

Download Full-text

Three generative, lexicalised models for statistical parsing

10.3115/976909.979620 ◽

1997 ◽

Cited By ~ 84

Author(s):

Michael Collins

Keyword(s):

Statistical Parsing

Download Full-text

On statistical parsing of French with supervised and semi-supervised strategies

10.3115/1705475.1705483 ◽

2009 ◽

Cited By ~ 3

Author(s):

Marie Candito ◽

Benoît Crabbé ◽

Djamé Seddah

Keyword(s):

Statistical Parsing

Download Full-text

Statistical Parsing of Tree Wrapping Grammars

10.18653/v1/2020.coling-main.595 ◽

2020 ◽

Author(s):

Tatiana Bladier ◽

Jakub Waszczuk ◽

Laura Kallmeyer

Keyword(s):

Statistical Parsing

Download Full-text

How relevant is linguistics to computational linguistics?

Linguistic Issues in Language Technology ◽

10.33011/lilt.v6i.1249 ◽

2011 ◽

Vol 6 ◽

Author(s):

Mark Johnson

Keyword(s):

Machine Learning ◽

Computational Linguistics ◽

Linguistic Theory ◽

Statistical Techniques ◽

Engineering Applications ◽

The Past ◽

Statistical Parsing ◽

Types Of Information ◽

The Relationship ◽

Scientific Fields

I start by explaining what I take computational linguistics to be, and discuss the relationship between its scientific side and its engineering applications. Statistical techniques have revolutionised many scientific fields in the past two decades, including computational linguistics. I describe the evolution of my own research in statistical parsing and how that lead me away from focusing on the details of any specific linguistic theory, and to concentrate instead on discovering which types of information (i.e., features) are important for specific linguistic processes, rather than on the details of exactly how this information should be formalised. I end by describing some of the ways that ideas from computational linguistics, statistics and machine learning may have an impact on linguistics in the future.

Download Full-text

Using model-theoretic semantic interpretation to guide statistical parsing and word recognition in a spoken language interface

10.3115/1075096.1075163 ◽

2003 ◽

Cited By ~ 1

Author(s):

William Schuler

Keyword(s):

Word Recognition ◽

Spoken Language ◽

Semantic Interpretation ◽

Statistical Parsing ◽

Theoretic Semantic ◽

Model Theoretic Semantic

Download Full-text

Fast Statistical Parsing with Parallel Multiple Context-Free Grammars

10.3115/v1/e14-1039 ◽

2014 ◽

Cited By ~ 1

Author(s):

Krasimir Angelov ◽

Peter Ljunglöf

Keyword(s):

Multiple Context ◽

Statistical Parsing ◽

Context Free ◽

Context Free Grammars

Download Full-text

A Statistical Parsing Framework for Sentiment Classification

Computational Linguistics ◽

10.1162/coli_a_00221 ◽

2015 ◽

Vol 41 (2) ◽

pp. 293-336 ◽

Cited By ~ 22

Author(s):

Li Dong ◽

Furu Wei ◽

Shujie Liu ◽

Ming Zhou ◽

Ke Xu

Keyword(s):

Sentiment Analysis ◽

Sentiment Classification ◽

Training Data ◽

Data Sets ◽

Syntactic Parsing ◽

Benchmark Data ◽

Sentence Level ◽

Statistical Parsing ◽

Parse Trees ◽

Context Free

We present a statistical parsing framework for sentence-level sentiment classification in this article. Unlike previous works that use syntactic parsing results for sentiment analysis, we develop a statistical parser to directly analyze the sentiment structure of a sentence. We show that complicated phenomena in sentiment analysis (e.g., negation, intensification, and contrast) can be handled the same way as simple and straightforward sentiment expressions in a unified and probabilistic way. We formulate the sentiment grammar upon Context-Free Grammars (CFGs), and provide a formal description of the sentiment parsing framework. We develop the parsing model to obtain possible sentiment parse trees for a sentence, from which the polarity model is proposed to derive the sentiment strength and polarity, and the ranking model is dedicated to selecting the best sentiment tree. We train the parser directly from examples of sentences annotated only with sentiment polarity labels but without any syntactic annotations or polarity annotations of constituents within sentences. Therefore we can obtain training data easily. In particular, we train a sentiment parser, s.parser, from a large amount of review sentences with users' ratings as rough sentiment polarity labels. Extensive experiments on existing benchmark data sets show significant improvements over baseline sentiment classification approaches.

Download Full-text

Statistical Parsing Joakim Nivre

Handbook of Natural Language Processing ◽

10.1201/9781420085938-20 ◽

2010 ◽

pp. 261-290

Keyword(s):

Statistical Parsing

Download Full-text