Skip to main content
Alibaba CloudCustomer feedback

A user has stopped using API-based LLMs after finding that Alibaba Cloud's local model, Qwen-3.8-27B, is capable of completing complex software refactoring tasks unsupervised.

What happened

Post: "Qwen-3.8-27B is good enough that I stopped using API"

Source

RedditSep 24, 2026By u/Training-Respect8066

r/LocalLLaMA

Qwen-3.8-27B is good enough that I stopped using API

upvotes
641
comments
266

Post

Highlighted: the lines this signal was extracted from

Many a praise have been sung on Qwen-3.8, but here is mine. Qwen-3.8 and I had a rocky start, because it thinks so much. Watching it working is painful, so you have to stop doing that. You have to let it work unsupervised. And that's okay, because it really is able to complete complex refactors on its own, making good decisions along the way. Not perfect, but hey, neither is API. The model quant is Q4_K_S, context is quantized to Q8_0, which seems to be okay, quality wise. I use the official Qwen. Briefly tried Swift-Qwen, which is indeed faster, but I found it getting trapped in loops, which is very rare in vanilla Qwen. I am using Qwen-3.8 in the Pi agent without MCP and with the minimum amount of tools. Bash is all you need, but I keep the read, write, and edit tools. The edit tool in Pi is the weakest link, the model often has to retry edits, because it messed up the indentation. I am waiting for someone to come up with a more fault-tolerant edit in Pi. Probably I have to make one myself some day. As a sandbox I use docker. My Pi agent is running on a Raspberry Pi, which seems fitting. On my hardware and where I live, 1M tokens cost 2.4 cent (input) and 70 cent (output) which is comparable to the cheapest providers on nano-gpt.com.

Also quoted as evidence

  • [comment u/Express_Quail_1493] Yes qwen3.8-27b is the first local model that i can run that i can turn away and let if do work without worrying that i will do something stupid.

reddit.com/r/LocalLLaMA/comments/1wp0z3i/qwen3827b_is_good_enough_tha...Read the full source

Comments on the post

5 of 266 comments
  • “It is good, but only for "smaller" projects. I notice that its sometimes a lot of hustle ("but wait", "Now lets check again"..) and so on. The last project i have done with Qwen-3.8-27B took about 10 hours - with Claude this would have been finished in 30 minutes or something. So yes - its good, but in terms of speed the frontiers are still worth it.”

    u/caenum145 points · Sep 24, 2026View

  • “Hear hear. I'm "GPU poor" (16GB) but I agree. I'm using ByteShape's IQ4_XS together with Raymond's KV cache streaming (context > 160k) and DFlash2 (at low context) and it's faster than any MoE I've used as well as good enough for me to use it as my only model for software dev, cybersec and IT adm. https://github.com/troed/llama.cpp-adaptive-kv-streaming Podman containers for sandbox using a se”

    u/tsangberg84 points · Sep 24, 2026View

  • “Yes qwen3.8-27b is the first local model that i can run that i can turn away and let if do work without worrying that i will do something stupid.”

    u/Express_Quail_149341 points · Sep 24, 2026View

  • “qwen3.8-27b-nvfp4-mtp has been outstanding for me (with Cline) when hosted in LM Studio with temp set to .1 and the "Reasoning Budget" to 1024. The latter definitely eliminated the annoying over thinking.”

    u/Great_Guidance_844819 points · Sep 24, 2026View

  • “at those level of quants, its pretty solid. You can minimize the overthinking to some degree without much a hit with SWIFT variants. The problem, as with many local models though, is context. Unless you're GPU rich, you're stuck in the 80k-120k at best context range, and thats with q4 context. Its painful, and you have to setup your harness right so it plans, and breaks the work into small chunk”

    u/ailee4312 points · Sep 24, 2026View

Extracted by Autobound

From the Signal API record
Signal
Customer feedback

What this signalsUser posts often show product pain before it reaches reviews or churn.

Subreddit
r/LocalLLaMA
Stage
Switched
Event date
Sep 2026

Companies

  • OpenAIAlso named
  • AnthropicAlso named
  • xAIAlso named

The full record

From the Signal API record

Numbers

Mentions
8

Details

Timing
Completed
Category
Features
Virality
Very high
Post kind
Text
Prominence
Core
Company's role
Vendor

Topics and mentions

Topics

  • local models
  • code generation
  • software development
  • ai
  • llm

Flair

  • Discussion

Products named

  • Qwen-3.8-27B

Extraction

Sentiment
Positive
Detected
Sep 24, 2026
signal_type
reddit-company
signal_subtype
customerFeedback

Use this data

Get every Reddit signal for Alibaba Cloud and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at Alibaba Cloud this week?”

  2. Send it to your own tools

    The Signal API returns Reddit signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full reddit-company record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/e8093e10-3f39-5f18-a3cf-7cf57da9a513 returns this record as JSON. POST /v1/companies/enrich returns every signal for alibabacloud.com.

{
  "signal_id": "e8093e10-3f39-5f18-a3cf-7cf57da9a513",
  "signal_type": "reddit-company",
  "signal_subtype": "customerFeedback",
  "detected_at": "2026-09-24T12:59:28+00:00",
  "company": {
    "name": "Alibaba Cloud",
    "domain": "alibabacloud.com"
  },
  "data": {
    "nsfw": false,
    "stage": "switched",
    "awards": 0,
    "timing": "completed",
    "topics": [
      "ai",
      "llm",
      "local models",
      "code generation",
      "software development"
    ],
    "post_id": "1wp0z3i",
    "summary": "A user has stopped using API-based LLMs after finding that Alibaba Cloud's local model, Qwen-3.8-27B, is capable of completing complex software refactoring tasks unsupervised.",
    "category": "features",
    "comments": [
      {
        "url": "https://www.reddit.com/r/LocalLLaMA/comments/1wp0z3i/comment/pbrf765/",
        "depth": 0,
        "score": 145,
        "author": "caenum",
        "excerpt": "It is good, but only for \"smaller\" projects. I notice that its sometimes a lot of hustle (\"but wait\", \"Now lets check again\"..) and so on. The last project i have done with Qwen-3.8-27B took about 10 hours - with Claude this would have been finished in 30 minutes or something. So yes - its good, but in terms of speed the frontiers are still worth it.",
        "posted_at": "2026-09-24T13:12:34.000Z",
        "author_url": "https://www.reddit.com/user/caenum/"
      },
      {
        "url": "https://www.reddit.com/r/LocalLLaMA/comments/1wp0z3i/comment/pbremm1/",
        "depth": 0,
        "score": 84,
        "author": "tsangberg",
        "excerpt": "Hear hear. I'm \"GPU poor\" (16GB) but I agree. I'm using ByteShape's IQ4_XS together with Raymond's KV cache streaming (context > 160k) and DFlash2 (at low context) and it's faster than any MoE I've used as well as good enough for me to use it as my only model for software dev, cybersec and IT adm.\n\n https://github.com/troed/llama.cpp-adaptive-kv-streaming\n\n Podman containers for sandbox using a se",
        "posted_at": "2026-09-24T13:09:43.000Z",
        "author_url": "https://www.reddit.com/user/tsangberg/"
      },
      {
        "url": "https://www.reddit.com/r/LocalLLaMA/comments/1wp0z3i/comment/pbresxo/",
        "depth": 0,
        "score": 41,
        "author": "Express_Quail_1493",
        "excerpt": "Yes qwen3.8-27b is the first local model that i can run that i can turn away and let if do work without worrying that i will do something stupid.",
        "posted_at": "2026-09-24T13:10:36.000Z",
        "author_url": "https://www.reddit.com/user/Express_Quail_1493/"
      },
      {
        "url": "https://www.reddit.com/r/LocalLLaMA/comments/1wp0z3i/comment/pbrd9b5/",
        "depth": 0,
        "score": 19,
        "author": "Great_Guidance_8448",
        "excerpt": "qwen3.8-27b-nvfp4-mtp has been outstanding for me (with Cline) when hosted in LM Studio with temp set to .1 and the \"Reasoning Budget\" to 1024. The latter definitely eliminated the annoying over thinking.",
        "posted_at": "2026-09-24T13:02:50.000Z",
        "author_url": "https://www.reddit.com/user/Great_Guidance_8448/"
      },
      {
        "url": "https://www.reddit.com/r/LocalLLaMA/comments/1wp0z3i/comment/pbriymd/",
        "depth": 0,
        "score": 12,
        "author": "ailee43",
        "excerpt": "at those level of quants, its pretty solid. You can minimize the overthinking to some degree without much a hit with SWIFT variants.\n\n The problem, as with many local models though, is context. Unless you're GPU rich, you're stuck in the 80k-120k at best context range, and thats with q4 context. Its painful, and you have to setup your harness right so it plans, and breaks the work into small chunk",
        "posted_at": "2026-09-24T13:30:45.000Z",
        "author_url": "https://www.reddit.com/user/ailee43/"
      },
      {
        "url": "https://www.reddit.com/r/LocalLLaMA/comments/1wp0z3i/comment/pbse3b9/",
        "depth": 0,
        "score": 7,
        "author": "relmny",
        "excerpt": "\"the model often has to retry edits, because it messed up the indentation\"\n\n that might be the result of quantizing KV and/or a low q4 quant. If you can try not quantizing KV and/or a higher quant, that might fix most of the issues.",
        "posted_at": "2026-09-24T15:47:48.000Z",
        "author_url": "https://www.reddit.com/user/relmny/"
      },
      {
        "url": "https://www.reddit.com/r/LocalLLaMA/comments/1wp0z3i/comment/pbrowux/",
        "depth": 0,
        "score": 6,
        "author": "bakatristan",
        "excerpt": "27b being this good is actually fucked lol. at this rate the only thing keeping me on API models is 24/7 uptime bc hosting it sometimes is a pain",
        "posted_at": "2026-09-24T13:58:23.000Z",
        "author_url": "https://www.reddit.com/user/bakatristan/"
      },
      {
        "url": "https://www.reddit.com/r/LocalLLaMA/comments/1wp0z3i/comment/pbsc5iz/",
        "depth": 0,
        "score": 6,
        "author": "MindfulMan1984",
        "excerpt": "Same, it passed the good enough on my use cases. Bye, OpenAI/Anthropic/xAI; you guys can \"slow down\" your \"frontier\" AI at any time. Soon everyone will realize that no one needs \"doomsday AI\" for everyday use.",
        "posted_at": "2026-09-24T15:39:43.000Z",
        "author_url": "https://www.reddit.com/user/MindfulMan1984/"
      }
    ],
    "evidence": [
      "[post] Qwen-3.8-27B is good enough that I stopped using API",
      "[post] it really is able to complete complex refactors on its own, making good decisions along the way. Not perfect, but hey, neither is API.",
      "[comment u/Express_Quail_1493] Yes qwen3.8-27b is the first local model that i can run that i can turn away and let if do work without worrying that i will do something stupid."
    ],
    "virality": "very_high",
    "post_date": "2026-09-24T12:59:28.000Z",
    "post_kind": "text",
    "post_text": "Many a praise have been sung on Qwen-3.8, but here is mine.\n\nQwen-3.8 and I had a rocky start, because it thinks so much. Watching it working is painful, so you have to stop doing that. You have to let it work unsupervised. And that's okay, because it really is able to complete complex refactors on its own, making good decisions along the way. Not perfect, but hey, neither is API.\n\nThe model quant is Q4_K_S, context is quantized to Q8_0, which seems to be okay, quality wise. I use the official Qwen. Briefly tried Swift-Qwen, which is indeed faster, but I found it getting trapped in loops, which is very rare in vanilla Qwen.\n\nI am using Qwen-3.8 in the Pi agent without MCP and with the minimum amount of tools. Bash is all you need, but I keep the read, write, and edit tools. The edit tool in Pi is the weakest link, the model often has to retry edits, because it messed up the indentation. I am waiting for someone to come up with a more fault-tolerant edit in Pi. Probably I have to make one myself some day.\n\nAs a sandbox I use docker. My Pi agent is running on a Raspberry Pi, which seems fitting.\n\nOn my hardware and where I live, 1M tokens cost 2.4 cent (input) and 70 cent (output) which is comparable to the cheapest providers on nano-gpt.com.",
    "sentiment": "positive",
    "subreddit": "LocalLLaMA",
    "event_date": "2026-09",
    "post_flair": [
      "Discussion"
    ],
    "post_title": "Qwen-3.8-27B is good enough that I stopped using API",
    "prominence": "core",
    "source_url": "https://www.reddit.com/r/LocalLLaMA/comments/1wp0z3i/qwen3827b_is_good_enough_that_i_stopped_using_api/",
    "entity_role": "vendor",
    "post_author": "Training-Respect8066",
    "upvote_ratio": 0.9637681159420289,
    "mention_count": 8,
    "mention_surge": true,
    "subreddit_url": "https://www.reddit.com/r/LocalLLaMA/",
    "total_upvotes": 641,
    "comments_total": 267,
    "total_comments": 266,
    "other_companies": [
      {
        "name": "OpenAI",
        "role": "replaced",
        "domain": "openai.com"
      },
      {
        "name": "Anthropic",
        "role": "replaced",
        "domain": "anthropic.com"
      },
      {
        "name": "xAI",
        "role": "replaced",
        "domain": "x.ai"
      }
    ],
    "post_author_url": "https://www.reddit.com/user/Training-Respect8066/",
    "signal_category": "feedback",
    "comments_included": 15,
    "products_mentioned": [
      "Qwen-3.8-27B"
    ]
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.