6.0 Module 6: Practical Implementation – Named Entity Recognition (NER)
6.1 Introduction: Identifying Real-World Entities in Text
Named Entity Recognition (NER) is the task of locating and classifying named entities in unstructured text into pre-defined categories. These entities can be the names of persons, organizations, locations, dates, monetary values, and more. NER is a cornerstone of information extraction and plays a vital role in a wide range of applications, including enhancing search algorithms, powering question-answering systems, and automatically constructing knowledge graphs from text corpora. It transforms a flat string of text into structured, meaningful information.
6.2 The OpenNLP NER Workflow
OpenNLP’s approach to NER is model-driven, relying on pre-trained statistical models to identify entities. For the English language, OpenNLP provides several specialized models, including:
- en-ner-person.bin (for person names)
- en-ner-location.bin (for locations)
- en-ner-organization.bin (for organizations)
- en-ner-date.bin (for dates)
- en-ner-time.bin (for times)
The workflow for performing NER is consistent regardless of the entity type:
- Load the appropriate NER model (e.g., en-ner-person.bin) using the TokenNameFinderModel class.
- Instantiate the NameFinderME class with the loaded model.
- Pass a tokenized sentence (as a String array) to the find() method, which returns the spans of any detected entities.
6.3 Practical Implementation: Finding Person Names in a Sentence
Let’s walk through an example of finding person names in a sentence.
import java.io.FileInputStream;
import java.io.InputStream;
import opennlp.tools.namefind.NameFinderME;
import opennlp.tools.namefind.TokenNameFinderModel;
import opennlp.tools.util.Span;
public class NameFinderME_Example {
public static void main(String args[]) throws Exception{
//Loading the NER – Person model
InputStream inputStream = new FileInputStream(“C:/OpenNLP_models/en-ner-person.bin”);
TokenNameFinderModel model = new TokenNameFinderModel(inputStream);
//Instantiating the NameFinder class
NameFinderME nameFinder = new NameFinderME(model);
//Getting the sentence in the form of String array
String [] sentence = new String[]{
“Mike”,
“and”,
“Smith”,
“are”,
“good”,
“friends”
};
//Finding the names in the sentence
Span nameSpans[] = nameFinder.find(sentence);
//Printing the spans of the names in the sentence
for(Span s: nameSpans)
System.out.println(s.toString());
}
}
Code Deconstruction:
- The en-ner-person.bin model is loaded into a TokenNameFinderModel object.
- A NameFinderME instance is created using this model.
- Crucially, the input is not a raw string but a String array, sentence, where each element is a pre-tokenized word. The NER model operates on tokens, not raw text.
- The nameFinder.find(sentence) method is called, and it returns an array of Span objects.
- The output, [0..1) person and [2..3) person, indicates that the model identified two entities of type “person”. The [start..end) notation here refers to the indices of the input token array, not character offsets. The first span identifies the token at index 0 (“Mike”), and the second identifies the token at index 2 (“Smith”).
A more complete and realistic example would first tokenize a raw sentence before passing it to the name finder.
import java.io.FileInputStream;
import java.io.InputStream;
import opennlp.tools.namefind.NameFinderME;
import opennlp.tools.namefind.TokenNameFinderModel;
import opennlp.tools.tokenize.TokenizerME;
import opennlp.tools.tokenize.TokenizerModel;
import opennlp.tools.util.Span;
public class NameFinderSentences {
public static void main(String args[]) throws Exception{
// Tokenize the sentence first
InputStream inputStreamTokenizer = new FileInputStream(“C:/OpenNLP_models/en-token.bin”);
TokenizerModel tokenModel = new TokenizerModel(inputStreamTokenizer);
TokenizerME tokenizer = new TokenizerME(tokenModel);
String sentence = “Mike is senior programming manager and Rama is a clerk both are working at Tutorialspoint”;
String tokens[] = tokenizer.tokenize(sentence);
// Load the NER model and find names
InputStream inputStreamNameFinder = new FileInputStream(“C:/OpenNLP_models/en-ner-person.bin”);
TokenNameFinderModel model = new TokenNameFinderModel(inputStreamNameFinder);
NameFinderME nameFinder = new NameFinderME(model);
Span nameSpans[] = nameFinder.find(tokens);
// Print the names and their spans
for(Span s: nameSpans)
System.out.println(s.toString()+” “+tokens[s.getStart()]);
}
}
This improved example correctly demonstrates the pipeline: tokenize the raw text first, then pass the resulting token array to the NameFinderME. The output loop then uses the span’s start index to retrieve the actual token text, producing a more readable result like [0..1) person Mike.
6.4 Extending NER: Detecting Different Entity Types
The power of OpenNLP’s NER framework lies in its modularity. To detect a different type of entity, such as a location, the process remains identical; you simply need to load the corresponding model.
The following example finds location names by loading en-ner-location.bin.
import java.io.FileInputStream;
import java.io.InputStream;
import opennlp.tools.namefind.NameFinderME;
import opennlp.tools.namefind.TokenNameFinderModel;
import opennlp.tools.tokenize.TokenizerME;
import opennlp.tools.tokenize.TokenizerModel;
import opennlp.tools.util.Span;
public class LocationFinder {
public static void main(String args[]) throws Exception{
InputStream inputStreamTokenizer = new FileInputStream(“C:/OpenNLP_models/en-token.bin”);
TokenizerModel tokenModel = new TokenizerModel(inputStreamTokenizer);
String paragraph = “Tutorialspoint is located in Hyderabad”;
TokenizerME tokenizer = new TokenizerME(tokenModel);
String tokens[] = tokenizer.tokenize(paragraph);
//Loading the NER-location model
InputStream inputStreamNameFinder = new FileInputStream(“C:/OpenNLP_models/en-ner-location.bin”);
TokenNameFinderModel model = new TokenNameFinderModel(inputStreamNameFinder);
NameFinderME nameFinder = new NameFinderME(model);
Span nameSpans[] = nameFinder.find(tokens);
for(Span s: nameSpans)
System.out.println(s.toString()+” “+tokens[s.getStart()]);
}
}
The key difference is the line InputStream inputStreamNameFinder = new FileInputStream(“C:/OpenNLP_models/en-ner-location.bin”);. With this change, the program correctly processes the sentence “Tutorialspoint is located in Hyderabad” and produces the output [4..5) location Hyderabad, demonstrating its ability to identify “Hyderabad” as a location.
6.5 Evaluating Model Confidence: Understanding NER Probabilities
The NameFinderME class provides a probs() method to evaluate the confidence of its predictions. After a call to find(), this method can be invoked to get the probabilities associated with the last sequence of entity tags that were decoded. This returns an array of doubles representing the model’s confidence for each tag assignment (e.g., person, not-a-person).
Although the provided source context contains a file named TokenizerMEProbs.java under the “NameFinder Probability” section, this appears to be a copy-paste error in the source. Conceptually, to get the NER probabilities, one would perform the following steps after finding names:
// … after calling nameFinder.find(tokens)
double[] probs = nameFinder.probs();
// Loop through the probabilities to analyze them
for (double p : probs) {
System.out.println(p);
}
This would provide a sequence of scores reflecting the model’s certainty about its classification for each token in the input sentence, which is invaluable for assessing the reliability of the extracted entities.
Now that we have explored how to identify what words are (entities), our next module will focus on identifying their grammatical function through Part-of-Speech tagging.