Update consuming_tgi.md

2025-09-10 20:04:52 +00:00 · 2023-08-18 17:14:52 +03:00 · 2023-08-18 17:14:52 +03:00 · 8f1d266e69
commit 8f1d266e69
parent 3a2a13ecd5
1 changed files with 2 additions and 6 deletions
--- a/docs/source/basic_tutorials/consuming_tgi.md
+++ b/docs/source/basic_tutorials/consuming_tgi.md
@ -143,13 +143,9 @@ You can try the demo directly here 👇

 You can disable streaming mode using `return` instead of `yield` in your inference function, like below.

-```diff
+```python
 def inference(message, history):
-    partial_message = ""
-    for token in client.text_generation(message, max_new_tokens=20, stream=True):
-        partial_message += token
-       yield partial_message
-+       return partial_message
+    return client.text_generation(message, max_new_tokens=20)
 ```

 You can read more about how to customize a `ChatInterface` [here](https://www.gradio.app/guides/creating-a-chatbot-fast).