FOSS · MIT—Open-source agent-swarm, the operating system for all your AI agents→
How-To/Build AI Agents from Scratch Locally: Step-by-Step | desplega.ai

Build AI Agents from Scratch Locally: Step-by-Step | desplega.ai

Learn how to build AI agents from scratch on your local machine—set up your environment, define tools, wire an LLM, and run your first autonomous agent loop.

Building an AI agent from scratch on your local machine is the fastest way to understand how modern autonomous systems actually work under the hood. Instead of relying on heavy frameworks that hide the reasoning loop, you'll wire the pieces together yourself: an LLM, a set of tools, a message history, and a control loop that decides when to keep going and when to stop. Once you understand this core pattern, extending it with memory, retrieval, or multi-agent orchestration becomes straightforward. If you plan to expose your agent through the Model Context Protocol later, our MCP best practices guide is a great follow-up.

Set Up Your Local Python Environment

Create a dedicated virtual environment and install the core dependencies your agent will need, including an LLM SDK and any tool libraries. Isolating dependencies keeps experiments reproducible and prevents version conflicts with other projects on your machine.

python -m venv .venv
source .venv/bin/activate  # Windows: .venv\Scripts\activate
pip install openai python-dotenv

Configure Your LLM API Credentials

Create a .env file at the project root and add your API key so the agent can authenticate against the language model provider without hardcoding secrets. Add .env to your .gitignore immediately to avoid leaking credentials into version control.

# .env
OPENAI_API_KEY=sk-...

Define Your Agent's Tools

Write plain Python functions that represent discrete actions your agent can take, then describe each tool in a JSON schema so the LLM knows when and how to call them. Good tools are small, deterministic, and well-named — the model uses the description and parameter names to decide when to invoke them, so treat those strings as prompts in their own right.

tools = [
  {
    "type": "function",
    "function": {
      "name": "search_web",
      "description": "Search the web and return a short summary.",
      "parameters": {
        "type": "object",
        "properties": {
          "query": {"type": "string", "description": "Search query"}
        },
        "required": ["query"]
      }
    }
  }
]

def search_web(query: str) -> str:
    # replace with real search logic
    return f"Results for: {query}"

Build the Agent Reasoning Loop

Implement a while loop that sends the conversation to the LLM, checks whether it requested a tool call, executes the matching function, and appends the result back into the message history. This tight loop — think, act, observe, repeat — is the beating heart of every agent framework you've ever seen. Everything else is convenience on top.

import json, os
from openai import OpenAI
from dotenv import load_dotenv

load_dotenv()
client = OpenAI()

def run_agent(user_input: str):
    messages = [{"role": "user", "content": user_input}]

    while True:
        response = client.chat.completions.create(
            model="gpt-4o",
            messages=messages,
            tools=tools,
            tool_choice="auto"
        )
        msg = response.choices[0].message
        messages.append(msg)

        if not msg.tool_calls:
            return msg.content  # final answer

        for tc in msg.tool_calls:
            fn_name = tc.function.name
            fn_args = json.loads(tc.function.arguments)
            result = globals()[fn_name](**fn_args)
            messages.append({
                "role": "tool",
                "tool_call_id": tc.id,
                "content": result
            })

Add a System Prompt to Shape Agent Behavior

Prepend a system message to the conversation that defines the agent's persona, constraints, and output format so responses remain consistent across runs. This is also where you enforce safety guardrails, tool-usage preferences, and citation requirements — all before the user ever types a word.

SYSTEM_PROMPT = """
You are a helpful research assistant. Always:
- Use the search_web tool before answering factual questions.
- Cite your sources.
- Reply in concise bullet points.
"""

def run_agent(user_input: str):
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user",   "content": user_input}
    ]
    # ... rest of loop

Run and Test Your Agent Locally

Execute the agent script from your terminal, pass a real query, and verify that the tool is called and a coherent final answer is returned before adding more complexity. Watch the message history print out step by step — this transparency is exactly why building from scratch beats black-box frameworks when you're still learning.

# agent.py — add at the bottom
if __name__ == "__main__":
    answer = run_agent("What is the latest stable version of Python?")
    print(answer)

# Then run:
# python agent.py

Write E2E Tests to Validate Agent Flows

Use desplega.ai to record and assert the full agent conversation flow — tool invocations, intermediate messages, and final output — so regressions are caught automatically on every code change. Agents are non-deterministic by nature, which makes traditional unit tests brittle. Flow-level assertions on latency, tool usage, and semantic output content are far more reliable signals of health.

# Example desplega test assertion (pseudo-code)
desplega.run_flow("agent-research-flow", inputs={"query": "Python latest version"})
  .assert_tool_called("search_web")
  .assert_final_output_contains("3.")
  .assert_latency_under_ms(5000)

From here, the natural next steps are adding persistent memory, plugging in retrieval over your own documents, and eventually exposing your tools over MCP so editors like Cursor and Claude Code can drive your agent directly. Keep the loop small, keep the tools sharp, and let your test suite tell you when something drifts.