SECAI Core · phase 1 of 15 · noun
Reinforcement learning
A type of machine learning that uses trial and error to make improved decisions by iterating through possible solutions to maximize rewards or avoid penalties.
The Explain card
- Plain English
- Reinforcement learning trains an agent by trial and error. It tries actions, collects rewards or penalties, and iterates until it finds a strategy that maximizes reward.
- Example
- An automated penetration testing agent is rewarded for each host it reaches in a lab network and penalized for tripping alarms. Over thousands of runs it learns quiet lateral movement paths.
- Why it matters
- The agent optimizes whatever reward you wrote, not what you meant. A badly specified reward produces clever, unexpected and sometimes unsafe behavior, and attackers can tamper with reward signals too.
- Hook
- Teach a dog with treats and it learns to get treats, not necessarily manners.
Word knowledge
How the word is built, where it came from, and what it sits beside in memory.
In a sentence
The pentest harness trains a reinforcement learning agent on a lab network and logs a reward after each hop that does not trip the IDS.
Why these words
- reinforcement French renforcer, to strengthen, from Latin fortis, strong supplies the strengthening half, so the learning hangs on a reward stamp rather than a labeled example
- learning Old English leornung, the act of getting knowledge names the activity the phrase assigns to an agent trained that way
Where it came from
- Origin
- English compound from French renforcer (to strengthen, from re-, again, and enforcer, to make strong, from Latin fortis, strong) and Old English leornung (the act of getting knowledge)
- Entered the language
- 1960s
- What changed
- Reinforce had meant to strengthen a wall or an army since about 1600. Learning had meant a person acquiring knowledge. Early behaviorists borrowed reinforcement for a stimulus that stamps in a response. Control labs in the 1960s joined the pair, and the compound became the ordinary name for agents trained against a score.
How it is spelled
- Pattern
- re- prefix meaning again or back, written solid with the stem
- Pattern
- -ment turns a verb into a noun of action or result
- Pattern
- force keeps Latin f, never ph
- Pattern
- -ing turns a verb into a noun of activity
- Breaks the pattern
- the noun phrase stays two words; a hyphen appears only when some styles treat it as a modifier
- Breaks the pattern
- the letters ei in reinforce are re meeting in, not the i-before-e vowel team
Spelled like
- enforcement
- payment
- training
Broken into chunks
-
re
- rewind
- rebuild
- retrain
-
force
- enforce
- forceful
- fortress
-
ment
- payment
- treatment
- movement
-
learn
- learner
- unlearn
- relearn
-
ing
- training
- computing
- reasoning
What it sits beside
Same subject
Same shape
- supervised learning
- unsupervised learning
- self-supervised learning
- federated learning
Learning compounds
Force family
- enforce
- enforcement
- fortress
Training regimes
- supervised learning
- unsupervised learning
- self-supervised learning
Where it sits in the deck
Phase 1: What AI Is: Core Concepts and Paradigms
You cannot secure, govern, or attack something you cannot define — establish what AI actually is before any other concept can land.