Natural language processing
From Wikipedia, the free
encyclopedia
Natural language processing (NLP)
is a field of computer science, artificial
intelligence, and linguistics concerned with the interactions
between computers and human (natural) languages. As such, NLP is related to the area of human–computer
interaction. Many challenges in NLP involve natural language
understanding -- that is, enabling computers to derive meaning from human
or natural language input.
History[edit]
The history of NLP
generally starts in the 1950s, although work can be found from earlier periods.
In 1950, Alan
Turing published an article titled "Computing Machinery and
Intelligence"
which proposed what is now called the Turing
test as a criterion of intelligence.
The Georgetown experiment in 1954 involved
fully automatic translation of more than sixty Russian sentences into English.
The authors claimed that within three or five years, machine translation would
be a solved problem.[2] However, real
progress was much slower, and after the ALPAC report in 1966, which found
that ten years long research had failed to fulfill the expectations, funding
for machine translation was dramatically reduced. Little further research in
machine translation was conducted until the late 1980s, when the first statistical machine
translation systems were
developed.
Some notably
successful NLP systems developed in the 1960s were SHRDLU, a natural language
system working in restricted "blocks
worlds"
with restricted vocabularies, and ELIZA, a simulation of a Rogerian psychotherapist, written by Joseph
Weizenbaumbetween
1964 to 1966. Using almost no information about human thought or emotion, ELIZA
sometimes provided a startlingly human-like interaction. When the
"patient" exceeded the very small knowledge base, ELIZA might provide
a generic response, for example, responding to "My head hurts" with
"Why do you say your head hurts?".
During the 1970s many
programmers began to write 'conceptual ontologies', which structured real-world
information into computer-understandable data. Examples are MARGIE (Schank,
1975), SAM (Cullingford, 1978), PAM (Wilensky, 1978), TaleSpin (Meehan, 1976),
QUALM (Lehnert, 1977), Politics (Carbonell, 1979), and Plot Units (Lehnert
1981). During this time, many chatterbots were written
including PARRY, Racter, and Jabberwacky.
Up to the 1980s, most
NLP systems were based on complex sets of hand-written rules. Starting in the
late 1980s, however, there was a revolution in NLP with the introduction of machine
learning algorithms for
language processing. This was due both to the steady increase in computational
power resulting from Moore's
Law and the gradual lessening of the
dominance of Chomskyan theories of
linguistics (e.g. transformational grammar), whose theoretical
underpinnings discouraged the sort of corpus linguistics that underlies the
machine-learning approach to language processing.[3] Some of the
earliest-used machine learning algorithms, such as decision
trees,
produced systems of hard if-then rules similar to existing hand-written rules.
Increasingly, however, research has focused onstatistical models, which make soft, probabilistic decisions based on
attaching real-valued weights to the features
making up the input data. The cache language models upon which many speech recognition systems now rely are
examples of such statistical models. Such models are generally more robust when
given unfamiliar input, especially input that contains errors (as is very
common for real-world data), and produce more reliable results when integrated
into a larger system comprising multiple subtasks.
Many of the notable
early successes occurred in the field of machine translation, due especially to work at IBM Research,
where successively more complicated statistical models were developed. These
systems were able to take advantage of existing multilingualtextual
corpora that had been produced
by the Parliament of Canada and the European
Union as a result of laws calling for the
translation of all governmental proceedings into all official languages of the
corresponding systems of government. However, most other systems depended on
corpora specifically developed for the tasks implemented by these systems,
which was (and often continues to be) a major limitation in the success of
these systems. As a result, a great deal of research has gone into methods of
more effectively learning from limited amounts of data.
Recent research has
increasingly focused on unsupervised and semi-supervised learning algorithms.
Such algorithms are able to learn from data that has not been hand-annotated
with the desired answers, or using a combination of annotated and non-annotated
data. Generally, this task is much more difficult than supervised learning, and typically produces less accurate
results for a given amount of input data. However, there is an enormous amount
of non-annotated data available (including, among other things, the entire
content of the World
Wide Web),
which can often make up for the inferior results.
Major tasks in NLP[edit]
The following is a
list of some of the most commonly researched tasks in NLP. Note that some of
these tasks have direct real-world applications, while others more commonly
serve as sub-tasks that are used to aid in solving larger tasks. What
distinguishes these tasks from other potential and actual NLP tasks is not only
the volume of research devoted to them but the fact that for each one there is
typically a well-defined problem setting, a standard metric for evaluating the
task, standard corpora on which the task can
be evaluated, and competitions devoted to the specific task.
Produce a readable
summary of a chunk of text. Often used to provide summaries of text of a known
type, such as articles in the financial section of a newspaper.
Given a sentence or
larger chunk of text, determine which words ("mentions") refer to the
same objects ("entities"). Anaphora resolution is a specific example
of this task, and is specifically concerned with matching up pronouns with the nouns or
names that they refer to. The more general task of coreference resolution also
includes identifying so-called "bridging relationships" involvingreferring expressions. For example, in a sentence such as "He
entered John's house through the front door", "the front door"
is a referring expression and the bridging relationship to be identified is the
fact that the door being referred to is the front door of John's house (rather
than of some other structure that might also be referred to).
This rubric includes
a number of related tasks. One task is identifying the discourse structure of
connected text, i.e. the nature of the discourse relationships between
sentences (e.g. elaboration, explanation, contrast). Another possible task is
recognizing and classifying the speech
acts in a chunk of text (e.g. yes-no
question, content question, statement, assertion, etc.).
Automatically
translate text from one human language to another. This is one of the most
difficult problems, and is a member of a class of problems colloquially termed
"AI-complete", i.e. requiring all of the different
types of knowledge that humans possess (grammar, semantics, facts about the
real world, etc.) in order to solve properly.
Separate words into
individual morphemes and identify the
class of the morphemes. The difficulty of this task depends greatly on the
complexity of the morphology (i.e. the structure
of words) of the language being considered. English has fairly simple
morphology, especially inflectional morphology, and thus it is often possible to
ignore this task entirely and simply model all possible forms of a word (e.g.
"open, opens, opened, opening") as separate words. In languages such
as Turkish, however, such an
approach is not possible, as each dictionary entry has thousands of possible
word forms.
Given a stream of
text, determine which items in the text map to proper names, such as people or
places, and what the type of each such name is (e.g. person, location,
organization). Note that, although capitalization can aid in
recognizing named entities in languages such as English, this information
cannot aid in determining the type of named entity, and in any case is often
inaccurate or insufficient. For example, the first word of a sentence is also
capitalized, and named entities often span several words, only some of which
are capitalized. Furthermore, many other languages in non-Western scripts (e.g. Chinese or Arabic) do not have any
capitalization at all, and even languages with capitalization may not
consistently use it to distinguish names. For example, Germancapitalizes all nouns, regardless of
whether they refer to names, and French and Spanish do not capitalize
names that serve asadjectives.
Convert information
from computer databases into readable human language.
Convert chunks of
text into more formal representations such as first-order
logic structures that are easier for computer programs to
manipulate. Natural language understanding involves the identification of the
intended semantic from the multiple possible semantics which can be derived
from a natural language expression which usually takes the form of organized
notations of natural languages concepts. Introduction and creation of language
metamodel and ontology are efficient however empirical solutions. An explicit
formalization of natural languages semantics without confusions with implicit
assumptions such as closed world assumption (CWA) vs. open world assumption, or
subjective Yes/No vs. objective True/False is expected for the construction of
a basis of semantics formalization.[4]
Given an image
representing printed text, determine the corresponding text.
Given a sentence,
determine the part
of speech for each word. Many
words, especially common ones, can serve as multiple parts
of speech.
For example, "book" can be a noun ("the book on
the table") or verb ("to book a
flight"); "set" can be a noun, verb oradjective; and "out"
can be any of at least five different parts of speech. Some languages have more
such ambiguity than others. Languages with little inflectional morphology, such as English are particularly
prone to such ambiguity. Chinese is prone to such
ambiguity because it is a tonal
language during verbalization.
Such inflection is not readily conveyed via the entities employed within the
orthography to convey intended meaning.
Determine the parse
tree (grammatical analysis) of a given
sentence. The grammar for natural
languages is ambiguous and typical sentences
have multiple possible analyses. In fact, perhaps surprisingly, for a typical
sentence there may be thousands of potential parses (most of which will seem
completely nonsensical to a human).
Given a human-language
question, determine its answer. Typical questions have a specific right answer
(such as "What is the capital of Canada?"), but sometimes open-ended
questions are also considered (such as "What is the meaning of
life?").
Given a chunk of
text, identify the relationships among named entities (e.g. who is the wife of
whom).
Given a chunk of
text, find the sentence boundaries. Sentence boundaries are often marked by periods or other punctuation
marks,
but these same characters can serve other purposes (e.g. marking abbreviations).
Extract subjective
information usually from a set of documents, often using online reviews to
determine "polarity" about specific objects. It is especially useful
for identifying trends of public opinion in the social media, for the purpose
of marketing.
Given a sound clip of
a person or people speaking, determine the textual representation of the
speech. This is the opposite of text
to speech and is one of the
extremely difficult problems colloquially termed "AI-complete" (see above).
In natural
speech there are hardly any pauses between
successive words, and thus speech segmentation is a necessary
subtask of speech recognition (see below). Note also that in most spoken
languages, the sounds representing successive letters blend into each other in
a process termed coarticulation, so the conversion
of the analog signal to discrete characters can be a very difficult process.
Given a sound clip of
a person or people speaking, separate it into words. A subtask of speech recognition and typically grouped
with it.
Given a chunk of
text, separate it into segments each of which is devoted to a topic, and
identify the topic of the segment.
Separate a chunk of
continuous text into separate words. For a language like English, this is fairly
trivial, since words are usually separated by spaces. However, some written
languages like Chinese, Japanese and Thai do not mark word
boundaries in such a fashion, and in those languages text segmentation is a
significant task requiring knowledge of the vocabulary and morphology of words in the
language.
Many words have more
than one meaning; we have to select the meaning which makes the most
sense in context. For this problem, we are typically given a list of words and
associated word senses, e.g. from a dictionary or from an online resource such
as WordNet.
In some cases, sets
of related tasks are grouped into subfields of NLP that are often considered
separately from NLP as a whole. Examples include:
This is concerned
with storing, searching and retrieving information. It is a separate field
within computer science (closer to databases), but IR relies on some NLP
methods (for example, stemming). Some current research and applications seek to
bridge the gap between IR and NLP.
This is concerned in
general with the extraction of semantic information from text. This covers
tasks such as named entity recognition, Coreference
resolution, relationship extraction, etc.
No comments:
Post a Comment