── Replace the mechanism with "people who work there," and the whole picture comes into view ──

Cast — 9 employees at AI Inc.

01
Mr. Prompt
Customer / brings the question
"Could you take care of this?"
02
Context Window Receptionist
Front desk / manages capacity
"We fit 128k tokens in here!"
03
Token Worker
Factory line / cuts text into pieces
"We deal in 'pieces,' not words"
04
Embedding Artisan
Numeric translator / words → numbers
"'Dog' as numbers — here you go"
05
Attention Manager
Priority / decides what matters
"Important here, ignore there!"
06
HQ Inference Staff
Inference engine / predicts next word
"Next word is... this one!"
07
RAG Librarian
External reference / runs lookups
"Let me check the external library"
08
Agent Supervisor
External tools / calls APIs & tools
"I can use calculators, APIs, anything"
09
Hallucination Rookie
Overconfident newbie / handle with care
"Probably... like this!" (no basis)

The story — a day at AI Inc., from "customer arrives" to "answer delivered"

1. Morning, a customer walks in

A customer arrives at the glass-walled office. Mr. Prompt. Today's request:

Mr. Prompt:"I'd like a 200-word summary of why pharmaceutical companies matter."

2. Capacity check at the front desk

At the entrance, the Context Window Receptionist greets him. His job is to confirm whether the customer's request, plus internal documents, fits within today's allowed capacity. A token is a "piece" of text — much finer than a word.

Receptionist:"Got it. The question is 25 tokens, internal docs still have room. English takes fewer pieces, though."

3. The factory line cuts text into pieces

The request flows immediately onto the factory line. Standing at the conveyor belt are the Token Workers. They split text not into "words" as we think of them, but into much finer character fragments. "Pha/rma/ceuti/cal/ com/pan/y..." — an odd-looking split, but that's how AI works.

4. Turning words into numbers

The pieces go to the Embedding Artisan. "'Pharma' as numbers... a 768-dimensional vector. Done!" Words become massive blocks of numbers as they flow to the next step. AI isn't reading text — it's always computing numbers.

5. Deciding what matters

Enter the Attention Manager. Instead of treating all information equally, he decides on the spot what to focus on.

Manager:"'Why they matter' is the key today! Strong weight on its relation to 'pharmaceutical company'!"

6. Predicting the next word, relentlessly

Now the HQ Inference Staff (in fact, hundreds of billions of them) spring into action. Drawing on knowledge absorbed from reading the world's texts, they predict the "next word." "Pha/rma/ ind/us/try/ ex/ists/ to / preserve / patient / trust..." — one piece at a time, the answer is born. This is AI's true nature: next-token prediction in action.

7. When something's unknown, run to the library

Mid-way, one staff member looks troubled. "Hmm, if they ask about the latest regulation, our internal knowledge might be too old..." That's when the RAG Librarian is called in.

Librarian:"Let me run to the external library!" — fetching latest info from the external database (the "library") and handing it to the staff. That's RAG (Retrieval-Augmented Generation).

8. Tools belong to the trusted supervisor

"Let me know if you need calculations," the Agent Supervisor chimes in. Calculators, search engines, email, external APIs — he can call on all of them. What turned "AI that responds" into "AI that actually does work" is him.

9. But everyone keeps an eye on the rookie in the corner

In the corner sits a young employee who attracts wary glances from the rest of the team. The Hallucination Rookie. He has a habit of confidently answering even when no one asked.

Rookie:"The most trusted pharma company in the world is Astellas!"
Coworkers:"Wait, on what basis?!"
Rookie:"Well, it just kinda felt right..."

That's the real nature of AI hallucination. Plausible-but-wrong outputs don't come from malice — they come from the structure of AI itself. Which is exactly why human verification is essential.

10. Delivering the finished answer

Finally, the Context Window Receptionist takes the completed answer and hands it back to Mr. Prompt. "Here you go!" — text streams onto the screen. That's a day at AI Inc.

Cast → real AI terminology

CharacterReal termIn one line
Mr. PromptPromptThe question or instruction itself
Context Window ReceptionistContext WindowMax tokens handled at once
Token WorkerTokenizationSplitting text into AI-handleable units
Embedding ArtisanEmbeddingConverting words to numeric vectors
Attention ManagerAttention mechanismDeciding what matters in context
HQ Inference StaffInference / next-token predictionPredicting the next word from learned probabilities
RAG LibrarianRAG (Retrieval-Augmented Generation)Pulling in external knowledge to answer
Agent SupervisorAI AgentAI that calls external tools to act
Hallucination RookieHallucinationPlausible but incorrect output
If you've read this far, you can probably picture, faintly now, the "giant office" that's working quietly behind your ChatGPT screen. Even when vendor documents say "the token limit is...", "RAG over internal documents...", "agent functionality..." — none of it should feel scary anymore. Just think back to what each employee does.

Vol.2 & Vol.3 (Available now)

Vol.2: Hands-on — building source-consistent documents with AI using Phrase Structure Grammar. Not just reading: ten steps where you and your AI build a verification prompt together.
Vol.3: "Email, Minutes, Summaries — 30 minutes a day with AI." Three practical flows that recover 22–33 hours per month.