Gateway latency diagnostics split slow AI agent requests into admission, queue wait, handler, and first-response phases so operators can fix the real bottleneck before changing models.
AI agent startup latency is often a first-reply problem, not only a model problem. OpenClaw 2026.6.6 shows where to cache, trace and defer startup work.