AI Streaming Arrives All at Once Behind NGINX
Trace incremental AI output from the application through the proxy to the browser, separating buffering from idle timeouts and rendering delays.
The essentials
- If the application emits incremental output but the browser receives one large response at the end, a layer between them may be buffering.
- Compare direct application output with proxied output before changing the model.
On this page
If the application emits incremental output but the browser receives one large response at the end, a layer between them may be buffering. Compare direct application output with proxied output before changing the model.
Establish where streaming stops
Use a harmless request and record timestamps at three points: application emission, client receipt and UI rendering. A stream can arrive correctly while the frontend waits to render it.
For a controlled fixture, have a development endpoint emit several small messages at known intervals. This isolates the transport from model loading and generation speed.
Compare the same endpoint directly and through the normal proxy. Keep protocol and response format consistent.
Inspect buffering and timeouts separately
NGINX documents response buffering and proxy timeouts in its HTTP proxy module reference. Buffering can delay delivery; an idle timeout can terminate a stream. Those symptoms need different fixes.
| Symptom | Next check |
|---|---|
| Direct stream incremental, proxy batches | Proxy buffering |
| Both paths batch | Application or client behavior |
| Stream cuts off after silence | Idle timeout or upstream failure |
| Bytes arrive, UI stays blank | Frontend parsing and rendering |
Other proxies, CDNs and hosting platforms may have additional limits. A local NGINX change cannot override every upstream constraint.
Make the smallest route-specific change
Apply streaming behavior to the intended endpoint rather than disabling buffering across unrelated downloads or pages. Verify content type and framing expected by the client.
If using server-sent events, preserve valid event boundaries. If using another streaming protocol, follow its framing rules instead of assuming every newline is a complete message.
Consider whether compression or an application middleware layer accumulates data. Check actual received bytes rather than inferring behavior from configuration alone.
Handle quiet periods
A long model-loading step can produce no output before generation begins. If the protocol permits keepalive messages, use them deliberately and ensure the client ignores them appropriately.
Increasing timeouts can be necessary for a legitimate workload, but it should not conceal a permanently stuck operation. Keep cancellation and an overall task deadline.
Verify errors and cancellation
Test a normal stream, a quiet period, an upstream failure and a user cancellation. Confirm the client distinguishes completion from an interrupted response.
Record the effective proxy configuration and the fixture timing results. Do not report a model-speed improvement when the change only made existing output visible sooner.
For AI agent applications, this distinction matters: visible streaming is presentation, while task completion and tool side effects need their own state.
Related troubleshooting
This guide draws on the linked documentation. Examples are illustrative unless explicitly identified as measured results.
Practical guides published by Lucivo, developed with AI assistance and references to official documentation. Examples are illustrative unless a guide explicitly documents a hands-on test. Check the linked sources for current product details.
Related articles
API Key Committed to Git: What to Do Next
API Timeout: Is It Safe to Retry?
API Works in curl but Fails in the Browser
The Weekly Breakdown
High signal AI & software stories.
Direct to your inbox. No hype.
Independent analysis of AI models, developer tools, and computing architectures. Delivered every Sunday morning. 100% free.