Learn spaCy as an industrial-strength NLP pipeline: Doc/Token/Span objects, linguistic annotations, named entities, rule-based matching, custom components, efficient batch processing and production packaging.
Enter your name. Complete at least 12 practice cases to unlock the Certificate of Participation.
spaCy turns raw text into a Doc, then passes that Doc through ordered pipeline components that add annotations.
Think in terms of one processed Doc carrying many layers of linguistic evidence.
import spacy
nlp = spacy.load('en_core_web_sm')
doc = nlp('Miami services improved this week.')
print([(t.text, t.pos_) for t in doc])Think in terms of one processed Doc carrying many layers of linguistic evidence.
spaCy preserves the original text while exposing tokens, character offsets, sentences and spans for precise downstream logic.
Preserve offsets when extracted information must be traced back to the original document.
doc = nlp('First request. Second request.')
print([sent.text for sent in doc.sents])
print(doc[0:2].text)Preserve offsets when extracted information must be traced back to the original document.
Trained spaCy pipelines can add part-of-speech tags, lemmas, dependency relations and other contextual annotations.
A blank language object tokenizes text; trained annotations require the appropriate pipeline components.
for token in nlp('The analyst reviewed the report.'):
print(token.text, token.lemma_, token.pos_, token.dep_)A blank language object tokenizes text; trained annotations require the appropriate pipeline components.
NER identifies spans such as organizations, people, places, dates or domain-specific labels learned by the model.
Do not assume a general-purpose NER model knows your organization’s custom vocabulary.
doc = nlp('Microsoft opened an office in Miami in 2025.')
print([(ent.text, ent.label_) for ent in doc.ents])Do not assume a general-purpose NER model knows your organization’s custom vocabulary.
Matcher and PhraseMatcher let you capture deterministic token or phrase patterns where explicit business rules are valuable.
Rule-based NLP is not obsolete; it is often the right control layer around statistical models.
from spacy.matcher import Matcher
matcher = Matcher(nlp.vocab)
matcher.add('SERVICE_DELAY', [[{'LOWER':'service'},{'LOWER':'delay'}]])
print(matcher(nlp('A service delay was reported.')))Rule-based NLP is not obsolete; it is often the right control layer around statistical models.
Custom components let you add organization-specific logic to the same processing pipeline and Doc object.
A custom component belongs in the pipeline only when its inputs, outputs and ordering are clear.
from spacy.language import Language
@Language.component('flag_urgent')
def flag_urgent(doc):
doc.user_data['urgent'] = 'urgent' in doc.text.lower()
return doc
nlp.add_pipe('flag_urgent', last=True)A custom component belongs in the pipeline only when its inputs, outputs and ordering are clear.
nlp.pipe processes streams of texts efficiently and lets you disable unnecessary components to reduce cost.
Optimization starts by avoiding work you do not need, not by guessing at hardware.
texts = ['First ticket', 'Second ticket', 'Third ticket']
for doc in nlp.pipe(texts, batch_size=64):
print(len(doc))Optimization starts by avoiding work you do not need, not by guessing at hardware.
Production NLP needs the trained pipeline, configuration, custom code and regression tests to travel together.
Your production artifact is the whole language pipeline, not just a model weight file.
nlp.to_disk('service_nlp')
loaded = spacy.load('service_nlp')
print(loaded.pipe_names)Your production artifact is the whole language pipeline, not just a model weight file.
Open each item only after answering it in your own words.
The tokenizer creates a Doc from the raw text.
A Span is a contiguous slice of a Doc, often used for entities, phrases or extracted regions.
When explicit token patterns provide auditable and deterministic business logic.
It processes many texts efficiently as a stream and supports batching.
The trained pipeline, its configuration, required custom components and validation expectations.
| Need | Use this training when… | Control |
|---|---|---|
| Fast baseline | You need a defensible first model and workflow | Validate independently |
| Custom behavior | You need to move below the high-level API | Add complexity only for a requirement |
| Production | Production | Version data, code, artifacts and monitoring |
The technical concepts and code patterns in this training follow official project documentation. Validate package versions and environment compatibility before production use.
“I can design spaCy NLP pipelines that combine statistical annotations, entities, deterministic rules, custom components and efficient batch processing.”
Complete at least 12 of the 24 practice cases (50%) and enter your name.