r/openrouter
Has anyone else had endlessly nonsense output from Open Inference? (GLM 5.3 Flash / DeepSeek V4.1 Flash)
- upvotes
- 17
- comments
- 9
Post
Hi all, I want to know whether anyone else has run into this, or whether it's just me. Yesterday I was using GLM 5.3 Flash and DeepSeek V4.1 Flash through OpenRouter in Pi. After a few turns, both started vomiting a seemingly never ending list of unrelated words, until i stopped them. Most of the tokens were billed as reasoning. It happened with two unrelated models, which seemed odd, so digging deeper i found that every request had gone to the same provider, Open Inference, which serves both models in fp4. Once I blocked Open Inference in my settings, the same model in the same session went straight back to normal (on Wafer and DekaLLM). I also noticed that, for GLM, it was also the most expensive provider by a long shot (and one of the slowest), which makes me question why i was routed to Open Inference in the first place (I have no particular routing config/setup). I don't know how Open Inference runs its deployments or exactly how OpenRouter decides where to route. It could just be a bad deployment or a bug, looking at the token volume (https://openrouter.ai/provider/open-inference) there's been a big drop after Sep 26th, so maybe it is something happening systematically. Has anyone experienced this? am I missing something obvious? but also: shouldn't there be something on openrouter gauging provider output quality? this could have been very expensive, both for the...
Keep reading with a free account
The rest of this post, and every signal for DeepSeek, is in your free account.
Extracted from these lines
[comment u/Brilliant-Hall1387] I recommend using DeepSeek and GLM directly from the source companies in China, great pricing and you know you get the real thing at really good prices. At least try it, top up an account with 10 USD each and evaluate the difference. DeepSeek offers 2 USD top up if you want to start really low.