Inference & RAG
Streaming
A way of showing a model's answer word by word as it is being generated, instead of waiting for the full response.
A way of showing a model's answer word by word as it is being generated, instead of waiting for the full response.