Python Data Science Library Mastery • Training 22
Article-Training • Production NLP Pipelines

spaCy

Turn Raw Text into Structured Linguistic Objects with Fast, Composable Pipelines

Learn spaCy as an industrial-strength NLP pipeline: Doc/Token/Span objects, linguistic annotations, named entities, rule-based matching, custom components, efficient batch processing and production packaging.

Raw Text → Tokenizer → Pipeline Components → Doc → Entities/Relations → Rules → Batch → Deploy
tickets • contracts • notes • resumes
↓
📄
🧩
🏷️
🔎
🧲
⚙️
🚚
📦
↓
structured Doc → reusable NLP pipeline
8modules
24interactive practices
50%certificate unlock
6market-ready skills
Your Learning Record

Make the practice count

Enter your name. Complete at least 12 practice cases to unlock the Certificate of Participation.

Practice progress0 / 24
MODULE 01
📄

spaCy Mental Model: Pipeline + Doc

spaCy turns raw text into a Doc, then passes that Doc through ordered pipeline components that add annotations.

👁️
See it this way

Think in terms of one processed Doc carrying many layers of linguistic evidence.

Core ideas

  • The tokenizer creates the Doc first
  • Pipeline components enrich the same Doc
  • Token, Span and Doc expose structured annotations
  • The pipeline configuration is part of the model artifact
Try this
import spacy
nlp = spacy.load('en_core_web_sm')
doc = nlp('Miami services improved this week.')
print([(t.text, t.pos_) for t in doc])
✅

Think in terms of one processed Doc carrying many layers of linguistic evidence.

Practice the decision, not just the syntax

Practice 1
Which statement best matches spaCy Mental Model: Pipeline + Doc?
Practice 2
What is a practical control in this module?
Practice 3
What should you remember before production use?
MODULE 02
🧩

Doc, Token and Span Objects

spaCy preserves the original text while exposing tokens, character offsets, sentences and spans for precise downstream logic.

👁️
See it this way

Preserve offsets when extracted information must be traced back to the original document.

Core ideas

  • Token is one unit inside a Doc
  • Span is a slice of a Doc
  • Offsets let you map predictions back to source text
  • Sentence boundaries support document-level workflows
Try this
doc = nlp('First request. Second request.')
print([sent.text for sent in doc.sents])
print(doc[0:2].text)
✅

Preserve offsets when extracted information must be traced back to the original document.

Practice the decision, not just the syntax

Practice 4
Which statement best matches Doc, Token and Span Objects?
Practice 5
What is a practical control in this module?
Practice 6
What should you remember before production use?
MODULE 03
🏷️

Linguistic Annotations

Trained spaCy pipelines can add part-of-speech tags, lemmas, dependency relations and other contextual annotations.

👁️
See it this way

A blank language object tokenizes text; trained annotations require the appropriate pipeline components.

Core ideas

  • POS labels grammatical role
  • Lemma gives a base form
  • Dependency parsing describes syntactic relationships
  • Annotations depend on the loaded trained pipeline
Try this
for token in nlp('The analyst reviewed the report.'):
    print(token.text, token.lemma_, token.pos_, token.dep_)
✅

A blank language object tokenizes text; trained annotations require the appropriate pipeline components.

Practice the decision, not just the syntax

Practice 7
Which statement best matches Linguistic Annotations?
Practice 8
What is a practical control in this module?
Practice 9
What should you remember before production use?
MODULE 04
🔎

Named Entity Recognition

NER identifies spans such as organizations, people, places, dates or domain-specific labels learned by the model.

👁️
See it this way

Do not assume a general-purpose NER model knows your organization’s custom vocabulary.

Core ideas

  • Entities are spans with labels
  • Entity quality depends on domain fit
  • Inspect false positives and missed entities
  • Rules and trained NER can complement each other
Try this
doc = nlp('Microsoft opened an office in Miami in 2025.')
print([(ent.text, ent.label_) for ent in doc.ents])
✅

Do not assume a general-purpose NER model knows your organization’s custom vocabulary.

Practice the decision, not just the syntax

Practice 10
Which statement best matches Named Entity Recognition?
Practice 11
What is a practical control in this module?
Practice 12
What should you remember before production use?
MODULE 05
🧲

Rule-Based Matching

Matcher and PhraseMatcher let you capture deterministic token or phrase patterns where explicit business rules are valuable.

👁️
See it this way

Rule-based NLP is not obsolete; it is often the right control layer around statistical models.

Core ideas

  • Matcher works with token patterns
  • PhraseMatcher is efficient for phrase lists
  • Rules are auditable and easy to test
  • Use rules for precision where patterns are known
Try this
from spacy.matcher import Matcher
matcher = Matcher(nlp.vocab)
matcher.add('SERVICE_DELAY', [[{'LOWER':'service'},{'LOWER':'delay'}]])
print(matcher(nlp('A service delay was reported.')))
✅

Rule-based NLP is not obsolete; it is often the right control layer around statistical models.

Practice the decision, not just the syntax

Practice 13
Which statement best matches Rule-Based Matching?
Practice 14
What is a practical control in this module?
Practice 15
What should you remember before production use?
MODULE 06
⚙️

Custom Pipeline Components

Custom components let you add organization-specific logic to the same processing pipeline and Doc object.

👁️
See it this way

A custom component belongs in the pipeline only when its inputs, outputs and ordering are clear.

Core ideas

  • Components execute in pipeline order
  • Custom extensions can store derived attributes
  • Keep components deterministic when possible
  • Test dependencies between components explicitly
Try this
from spacy.language import Language
@Language.component('flag_urgent')
def flag_urgent(doc):
    doc.user_data['urgent'] = 'urgent' in doc.text.lower()
    return doc
nlp.add_pipe('flag_urgent', last=True)
✅

A custom component belongs in the pipeline only when its inputs, outputs and ordering are clear.

Practice the decision, not just the syntax

Practice 16
Which statement best matches Custom Pipeline Components?
Practice 17
What is a practical control in this module?
Practice 18
What should you remember before production use?
MODULE 07
🚚

Efficient Batch Processing with nlp.pipe

nlp.pipe processes streams of texts efficiently and lets you disable unnecessary components to reduce cost.

👁️
See it this way

Optimization starts by avoiding work you do not need, not by guessing at hardware.

Core ideas

  • Use nlp.pipe for many documents
  • Batching improves throughput
  • Disable unused components when safe
  • Measure throughput on representative text lengths
Try this
texts = ['First ticket', 'Second ticket', 'Third ticket']
for doc in nlp.pipe(texts, batch_size=64):
    print(len(doc))
✅

Optimization starts by avoiding work you do not need, not by guessing at hardware.

Practice the decision, not just the syntax

Practice 19
Which statement best matches Efficient Batch Processing with nlp.pipe?
Practice 20
What is a practical control in this module?
Practice 21
What should you remember before production use?
MODULE 08
📦

Package, Validate and Operate

Production NLP needs the trained pipeline, configuration, custom code and regression tests to travel together.

👁️
See it this way

Your production artifact is the whole language pipeline, not just a model weight file.

Core ideas

  • Record the exact pipeline package and version
  • Use representative regression examples
  • Track entity and rule coverage over time
  • Package custom components with deployment code
Try this
nlp.to_disk('service_nlp')
loaded = spacy.load('service_nlp')
print(loaded.pipe_names)
✅

Your production artifact is the whole language pipeline, not just a model weight file.

Practice the decision, not just the syntax

Practice 22
Which statement best matches Package, Validate and Operate?
Practice 23
What is a practical control in this module?
Practice 24
What should you remember before production use?
5-Question Knowledge Check

Can you explain the workflow before you write the code?

Open each item only after answering it in your own words.

1. What happens before spaCy pipeline components run?

The tokenizer creates a Doc from the raw text.

2. What is a Span?

A Span is a contiguous slice of a Doc, often used for entities, phrases or extracted regions.

3. When should you use Matcher?

When explicit token patterns provide auditable and deterministic business logic.

4. Why use nlp.pipe?

It processes many texts efficiently as a stream and supports batching.

5. What must move to production with a spaCy pipeline?

The trained pipeline, its configuration, required custom components and validation expectations.

Decision Guide

Where does spaCy fit?

NeedUse this training when…Control
Fast baselineYou need a defensible first model and workflowValidate independently
Custom behaviorYou need to move below the high-level APIAdd complexity only for a requirement
ProductionProductionVersion data, code, artifacts and monitoring
Official Sources & Further Learning

Grounded in the official spaCy documentation

The technical concepts and code patterns in this training follow official project documentation. Validate package versions and environment compatibility before production use.

Market Skills

What you should be able to say after this training

“I can design spaCy NLP pipelines that combine statistical annotations, entities, deterministic rules, custom components and efficient batch processing.”

Certificate of Participation

Unlock at 50% participation

Complete at least 12 of the 24 practice cases (50%) and enter your name.

0 / 24 • 0%