Imagine you walk into a library.
- Librarian 1 answers your question from what they already know.
- Librarian 2 can also search the internet and open your documents, finding things that weren't in their original knowledge.
- Librarian 3 takes a bigger job: "Research this topic, compare the options, prepare a report, and send it to my team." They break it into steps, use tools, check results, and keep going until it's done.
All three look like helpful assistants, but they work very differently. AI has the same three levels: chatbots that answer, applications that look things up, and agents that complete tasks. Anyone learning AI hits a wall of terms: LLMs, parameters, tokens, RAG, embeddings, MCP. Few explain how they fit together. So let's start at the beginning.
1. What Is an LLM?
LLM stands for Large Language Model. Picture a student who has read an enormous amount of books, articles, and documentation, and learned how words relate and how questions are usually answered.
An LLM is trained on huge amounts of text and learns patterns in language. When you ask something, it generates the answer one small piece at a time, each time predicting what should come next. These pieces are called tokens: a word, part of a word, or punctuation.
After this first stage, models usually get further training to be helpful and safe in conversation. That's why an assistant feels like a conversation partner and not just autocomplete.
2. What Are Parameters, and Why Do People Talk About Billions?
"7 billion parameters." "70 billion parameters." It sounds like phone specs, where bigger must be better. Not necessarily.
Parameters are the numerical values learned during training that shape the model's behaviour. A 70B model has about ten times as many as a 7B one. But training data, design, methods, and the task all matter. A smaller, well-built model can beat a larger one on specific jobs, while a bigger one costs more memory and computing power.
| Size | What it generally means |
|---|---|
| 1B–3B | Small enough for some laptops and phones |
| 7B–8B | Common size for capable smaller models |
| 30B–70B | Heavier to run, often stronger on complex tasks |
| Hundreds of billions | Needs serious infrastructure |
Many companies don't disclose model sizes, and designs like Mixture of Experts activate only part of the model per token, so even total size can mislead. The better question: "Better at what, and at what cost?"
3. Which LLMs Exist?
- OpenAI: GPT models, used in ChatGPT
- Google: Gemini
- Anthropic: Claude
- Meta: Llama
- DeepSeek, Alibaba (Qwen), Mistral: other major families
This list changes fast. People use "open-source AI" loosely for models whose weights you can download; more precisely these are "open-weight" models, and each licence decides what you may do with it. Even a free model costs money to run: hardware, electricity, or a cloud service.
4. What Is Temperature?
For a technical explanation you want a predictable answer. For a story about a robot opening a bakery on Mars, you may want surprising ideas. Temperature controls this: lower means more predictable word choices, higher allows more variety. It is not a measure of intelligence or accuracy. A low-temperature answer can still be wrong.
- Maximum output tokens: limit on response length
- Top-p: another way to restrict which words can be picked
- System instructions: set the model's role and rules
- Context window: how much the model can handle at once
These change how a model is used. They don't turn a weak model into a strong one.
5. What Is a Context Window? Is It Memory?
Imagine walking into a meeting with a huge folder, but the table only fits so much. A context window is that table: the information the model can look at in one request, including your question, earlier conversation, documents, and tool results. It's measured in tokens, not words.
- Models don't always use very long inputs equally well, and longer inputs cost more and run slower.
- A context window is not permanent memory. Remembering you across conversations is a separate feature an application must build.
6. What Is RAG?
Your company has 500 pages of internal documents. You ask an AI, "What's our leave policy?" The model knows policies in general but can't know yours unless someone provides it.
RAG (Retrieval-Augmented Generation) fixes this: the application first finds relevant information, then hands it to the model.
- You ask a question.
- The application searches your documents.
- It pulls out the most relevant passages.
- It gives them to the model.
- The model answers using your question and those passages.
RAG doesn't guarantee correctness. It can retrieve the wrong passage, documents may be outdated, or the model may misread them. Good systems track sources and are tested carefully.
7. What Are Embeddings?
"How do I apply for leave?" and "What's the procedure for requesting time off?" mean the same thing with different words. Keyword search might miss the match.
An embedding turns text (or an image) into a long list of numbers called a vector. Similar meanings get similar numbers. Imagine a giant map where related topics sit close together; your question is placed on the map and the system looks at what's nearby. This is semantic search. Tip: use the same embedding model for documents and questions, or the map won't line up.
A vector database stores embeddings and searches them quickly. The typical flow:
Documents → chunks → embeddings → vector storage → retrieval → LLM answer
- Chunking: splitting long documents into sections
- Metadata: extra info like document name or department
- Retrieval: finding passages relevant to a question
A vector database doesn't write answers, and RAG doesn't always need one. Smaller projects can use keyword search, regular databases, or a mix.
8. What Is an AI Agent?
A chatbot answers. A RAG app looks things up first. An agent uses a model to decide which actions to take toward a goal, uses tools, looks at results, and decides what to do next.
Ask: "Find five suitable companies, compare their openings with my skills, and prepare a shortlist."
- A chatbot explains how to search for jobs.
- A RAG app searches job descriptions you gave it.
- An agent searches career pages, extracts listings, compares them to your profile, ranks them, builds the shortlist, and asks approval before anything consequential, like applying.
Researching and acting are different, and good systems put approval steps around actions that matter. Agents usually combine a model, instructions, tools, memory or state, a loop that checks results, and rules limiting permissions and cost.
9. What Are Tools and Agent Frameworks?
An LLM can write about sending an email, but that doesn't mean it can send one. The application must provide a tool (like an email service) and let the model request it. The model asks; the application executes and returns the result. Examples: web search, database queries, calendars, file readers, code runners. This is the difference between describing an action and performing one.
Agent frameworks save developers from writing everything from scratch: model calls, state tracking, retries, workflow logic. Examples: LangChain, LangGraph, Semantic Kernel, LlamaIndex, OpenAI Agents SDK. Simple agents can be built without any framework. A framework isn't the agent's brain, just scaffolding.
10. What Is MCP?
Imagine every AI app needing a custom connection to every tool. The Model Context Protocol (MCP) is an open standard for connecting AI applications to external tools and data in a common way. An MCP server can offer tools (actions), resources (information), and prompts (reusable templates).
11. Other Terms You'll Keep Hearing
- Inference: using a trained model to produce an answer
- Training: adjusting parameters using data
- Fine-tuning: further training on selected examples
- Quantization: lower-precision values to save memory, sometimes slightly reducing quality
- Multimodal: handles images, audio, or video, not just text
- Hallucination: a confident answer that is wrong or made up
- API: a way for software to talk to other software
12. How It All Fits Together
An engineer asks: "Machine 24 stopped sending telemetry. Check the documentation and recent logs, find likely causes, and prepare a troubleshooting report."
- The application receives the request and passes it to an LLM.
- A RAG pipeline searches technical documents for similar failures.
- A tool pulls recent logs from an approved source.
- The model analyses documents and logs and drafts findings.
- If needed, the agent runs another check.
- The application produces a report with evidence, causes, and next steps.
- Anything affecting a live machine requires human approval and strict permissions.
13. Do You Need to Learn Everything?
No. Start with the problem you want to solve.
- Basic chatbot? Learn to call an LLM API.
- Answers from your documents? Learn chunking, embeddings, RAG.
- Live data or actions? Learn tool calling.
- Multistep tasks? Explore agent workflows.
- Running models locally? Study sizes, memory, licences.
- Standard tool connections? Learn MCP.
A sensible path: LLM basics, app development, RAG, tool use, agent workflows, then production concerns like testing, security, cost, and reliability. The best way to learn is to build something small: a document assistant that answers from your notes and shows its sources. Then add a tool. Only then ask whether you need an agent at all.
14. A Reality Check
- An LLM can be confidently wrong.
- RAG can retrieve outdated or irrelevant documents.
- Agents can pick the wrong tool or loop.
- A bigger model can cost more without being better for your task.
- A demo that works can fail on messy real-world data.
Good AI work means evaluating accuracy, speed, cost, security, and failure cases, and knowing when a simple database query beats an LLM or a human should check the result. The goal isn't AI everywhere. It's the right tool for a real problem.
The Bigger Picture
- Librarian 1 is the LLM, answering from what it learned.
- Librarian 2 is retrieval, bringing in outside information.
- Librarian 3 is the agent, using tools and working through steps.
Parameters shape behaviour. Context windows set how much the model sees. Embeddings and vector databases help find the right information. Frameworks coordinate the parts, and MCP standardises connections to outside tools. They aren't unrelated technologies, but pieces of one system.
Don't just memorise names. Learn why each piece exists, when it's useful, and what can go wrong. The future of AI belongs to people who can turn these building blocks into reliable solutions for real problems.


