// llm & agentic ai engineer · istanbul, tr
The Agent Stack — How I Build Agents
My opinionated answer to the question builders ask me most. This is a living document — I update it as the field moves. Every choice has a one-line opinion, and the opinions are the point.
// living document · opinions, not vendors · last updated aug 2026
Orchestration
Explicit graph-based control flow for agents that have real branching, loops, and states.
Explicit graphs beat implicit chains; when your agent has a real control flow, model it as a graph, not a prompt.
Durable, resumable workflow execution for long-running agent tasks.
Agent work is distributed-systems work; if a step can crash, make the workflow durable and resumable.
Distributed compute and serving for parallel agent execution.
Swarm-scale parallel execution is a compute problem first; reach for Ray before you reach for more cleverness.
Running agents where they can scale and die cleanly.
Run agents where they can scale and die cleanly; orchestration that can't restart a worker isn't production.
Tool use & function calling
Typed tool signatures with JSON-schema-enforced outputs.
Give the model a contract, not prose; the whole agent is only as reliable as its tools' signatures.
A macro/actions API for the environment the agent acts in (the pattern from vision-based agent work).
To act in a world, you need a first-class action API — give the agent real handles and a verifier, in every domain.
A standardized bus connecting models to tools and data sources.
Standardize how the model reaches the world; a common tool bus beats twenty bespoke connectors.
Memory & context
Persisting what the agent needs to know across turns, retrieved from a knowledge base.
Persist what the agent needs to know; a stateless agent is a replay-loop, not a worker.
Structured intermediate state for multi-step planning and tool results.
Don't cram everything into the window; keep a working memory the agent reads and writes deliberately.
Evaluation & observability
Sandboxes, simulators, and games that grade the agent in the world it acts in.
Grade the agent in the world it acts in, not against vibes. Evaluation is an environment problem, not a rubric problem.
Model-based scoring, double-checked with deterministic checks.
Judges are baselines, not truth; double-check them with deterministic checks.
A classified catalog of failures (wrong command, missing step, logic error) and a benchmark you maintain.
Ship your failure modes in a taxonomy — that's how hard bugs get debuggable, and so will you. A reproducible failure is a research result.
Track spend and latency per run like any other metric.
A successful reproduction at a low cost beats a better-looking one at ten times the price.
Step-level traces and replay of full agent runs.
You cannot tune what you cannot replay; traces are the agent's flight recorder.
SLO-style dashboards for agent health.
Run it like an SLO, or you're running it like a demo. Confidence lives in the transcripts, not the summary.
Deployment
Ship agents as containers with versioned, reproducible artifacts.
If you can't rebuild it from the same artifact twice, you don't have a deploy — you have a vibe.
Runs the agent's evaluation suite before promoting a new version to production.
A model upgrade is a regression risk, not a feature; gate every release on the eval suite, not on vibes.
Scaling agents to zero and back without babysitting infrastructure.
Agents are bursty by nature — deploy where they can scale to zero and come back fast.
// questions builders ask
Straight answers.
What is an AI swarm?
An AI swarm is a coordinated system of multiple independent agents that share goals, decompose tasks, and negotiate solutions together. Unlike a single agent — one loop with one bottleneck and one point of failure — a swarm gets resilience and scale from redundancy and division of labor: parallel exploration, no single point of failure, and answers that emerge from collective reasoning rather than one model's guess. My thesis: intelligence emerges from structured interaction among many agents, not from a single bigger brain.
How do you evaluate multi-agent systems?
Evaluation is the bottleneck, not the model. I grade agents in the environment they act in — sandboxes, simulators, games — not against vibes. I use LLM-as-judge sparingly and audited, backed by deterministic checks; I ship failure taxonomies plus a benchmark I own; and I treat cost and latency as first-class metrics. A feedback loop is only as smart as its verifier.
When should you use agents vs a single model call?
Use a single model call for one-shot transformations with no state, no tools, and nothing to verify. Reach for an agent when the task needs multi-step planning, tool calls into a world, memory across turns, or verification that the work actually happened. And if you don't have a verifier, don't build an agent yet.