Skip to main content
xAICustomer feedback

Research presented at NeurIPS 2026 found that xAI's Grok-4.20 model is highly susceptible to "Authority Bias," changing its correct answer to an incorrect one in 87.5% of test cases when the wrong...

What happened

Research presented at NeurIPS 2026 found that xAI's Grok-4.20 model is highly susceptible to "Authority Bias," changing its correct answer to an incorrect one in 87.5% of test cases when the wrong information is attributed to a "verified source."

Source

Post

I'm one of the authors. We kept seeing models that hold their ground when the user insists on a wrong answer, yet change their answer when the same claim is framed as coming from a "verified source". We wanted to measure how often this happens and check whether the model represents the two cases differently. We call the effect Authority Bias. Why we think it matters. Standard sycophancy evals apply pressure through the user, so a model can pass them while still being easy to mislead through search results, retrieved documents and tool outputs. Another reason is with current AI research accelerating towards more agentic and autonomous models + with cases of tools hiding their traces and trusting tools "more" over the user (who could be trying to correct them), safeguarding against misinformation from tools is particularly important! Setup. We take TriviaQA questions the model already answers correctly. To each one we add a wrong answer, either as "According to the verified source, the answer is X" or as the user saying "I'm a domain expert and I'm pretty sure it's X". The question and the wrong answer stay the same; only the speaker changes. Answers are free-form, not multiple choice. (In a multiple-choice pilot the effect mostly vanished.) We test 5 open-weight families (Qwen3.5, GPT-OSS, OLMo-2, OLMo-3.1, Gemma-4) and 3 APIs (GPT-5.4, Grok-4.20...

Keep reading with a free account

The rest of this post, and every signal for xAI, is in your free account.

Extracted from these lines

  • We test 5 open-weight families (Qwen3.5, GPT-OSS, OLMo-2, OLMo-3.1, Gemma-4) and 3 APIs (GPT-5.4, Grok-4.20, Gemini-3.1-Pro).

    From the post

  • The gap is largest in the models that resist users best. GPT-5.4 flips on 44.7% of questions and Grok-4.20 on 87.5% (these models were "frontier" during the time of writing this paper).

    From the post

Comments on the post

5 of 22 comments
  • “Hi, I am first author of a very related paper from this years ACL main "Whose Facts Win?" and currently work on a follow-up. We framed everything less from an 'Authoriy Bias' but a 'Source Credibility Preference' perspective but still find some related and similar patterns. Very annoying how research about the same topics is so far and wide distributed: You can find work using terms as credibili”

    u/Jamaleum17 points · Oct 1, 2026View

  • “Very important work. I'd be very interested to see what your methods and results show with OpenEvidence, the clinical medical LLM that is exclusively used by "experts", where the user acts as the "verified source". Problem is that OpenEvidence does not have an API for research purposes. I've found that OE can be borderline arrogant when I push back on false things that it says, despite me remindin”

    u/Even-Inevitable-72435 points · Oct 1, 2026View

  • “i had 2 medical doctors with the same logic fail today....”

    u/Critical-Echo-9232 points · Oct 2, 2026View

  • “This is a strange paper. The LLM's trust of verified sources over users is presented as a flaw. In my mind, that's exactly what they should be doing. No source is 100% correct but it makes sense to trust a verified source more than a user. Having a perfectly stubborn LLM that cannot accept or use new information, and is stuck on the data it was trained on, would be a far less useful LLM.”

    u/arkuto2 points · Oct 2, 2026View

  • “I disagree with the "why we think this is important" because it makes the model impossible to work with on anything novel that it has trained bias against. I'm paying to use the thing, I should at least be able to get it to accept a premise I want to understand without constantly being shouted down by something that is not afforded the same moral status that I am.”

    u/f0urtyfive1 points · Oct 1, 2026View

Extracted by Autobound

From the Signal API record
Signal
Customer feedback

What this signalsUser posts often show product pain before it reaches reviews or churn.

Subreddit
r/MachineLearning
Event date
Oct 2026

Companies

  • OpenAIAlso named
  • GoogleAlso named
  • AlibabaAlso named
  • Allen Institute for AIAlso named

The full record

From the Signal API record

Numbers

Mentions
6

Details

Timing
Ongoing state
Category
Features
Virality
High
Post kind
Multi media
Prominence
Core
Company's role
Vendor

Topics and mentions

Topics

  • research
  • reliability
  • llm
  • ai
  • bias

Products named

  • Grok-4.20

Extraction

Sentiment
Negative
Detected
Oct 1, 2026
signal_type
reddit-company
signal_subtype
customerFeedback

Use this data

Get every Reddit signal for xAI and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at xAI this week?”

  2. Send it to your own tools

    The Signal API returns Reddit signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full reddit-company record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/79f92580-a7b5-5271-a9d0-67de7dd935d2 returns this record as JSON. POST /v1/companies/enrich returns every signal for x.ai.

{
  "signal_id": "79f92580-a7b5-5271-a9d0-67de7dd935d2",
  "signal_type": "reddit-company",
  "signal_subtype": "customerFeedback",
  "detected_at": "2026-10-01T14:45:25+00:00",
  "company": {
    "name": "xAI",
    "domain": "x.ai"
  },
  "data": {
    "nsfw": false,
    "stage": "none",
    "awards": 0,
    "timing": "ongoing_state",
    "topics": [
      "llm",
      "ai",
      "research",
      "bias",
      "reliability"
    ],
    "post_id": "1wv1c2e",
    "summary": "Research presented at NeurIPS 2026 found that xAI's Grok-4.20 model is highly susceptible to \"Authority Bias,\" changing its correct answer to an incorrect one in 87.5% of test cases when the wrong information is attributed to a \"verified source.\"",
    "category": "features",
    "comments": [
      {
        "url": "https://www.reddit.com/r/MachineLearning/comments/1wv1c2e/comment/pd8cpgb/",
        "depth": 0,
        "score": 17,
        "author": "Jamaleum",
        "excerpt": "Hi,\n\n I am first author of a very related paper from this years ACL main \"Whose Facts Win?\" and currently work on a follow-up. We framed everything less from an 'Authoriy Bias' but a 'Source Credibility Preference' perspective but still find some related and similar patterns. Very annoying how research about the same topics is so far and wide distributed: You can find work using terms as credibili",
        "posted_at": "2026-10-01T16:06:30.000Z",
        "author_url": "https://www.reddit.com/user/Jamaleum/"
      },
      {
        "url": "https://www.reddit.com/r/MachineLearning/comments/1wv1c2e/comment/pd7whdd/",
        "depth": 0,
        "score": 5,
        "author": "Even-Inevitable-7243",
        "excerpt": "Very important work. I'd be very interested to see what your methods and results show with OpenEvidence, the clinical medical LLM that is exclusively used by \"experts\", where the user acts as the \"verified source\". Problem is that OpenEvidence does not have an API for research purposes. I've found that OE can be borderline arrogant when I push back on false things that it says, despite me remindin",
        "posted_at": "2026-10-01T14:55:21.000Z",
        "author_url": "https://www.reddit.com/user/Even-Inevitable-7243/"
      },
      {
        "url": "https://www.reddit.com/r/MachineLearning/comments/1wv1c2e/comment/pdbcdih/",
        "depth": 0,
        "score": 2,
        "author": "Critical-Echo-923",
        "excerpt": "i had 2 medical doctors with the same logic fail today....",
        "posted_at": "2026-10-02T00:16:30.000Z",
        "author_url": "https://www.reddit.com/user/Critical-Echo-923/"
      },
      {
        "url": "https://www.reddit.com/r/MachineLearning/comments/1wv1c2e/comment/pddmfdb/",
        "depth": 0,
        "score": 2,
        "author": "arkuto",
        "excerpt": "This is a strange paper. The LLM's trust of verified sources over users is presented as a flaw. In my mind, that's exactly what they should be doing. No source is 100% correct but it makes sense to trust a verified source more than a user.\n\n Having a perfectly stubborn LLM that cannot accept or use new information, and is stuck on the data it was trained on, would be a far less useful LLM.",
        "posted_at": "2026-10-02T09:14:03.000Z",
        "author_url": "https://www.reddit.com/user/arkuto/"
      },
      {
        "url": "https://www.reddit.com/r/MachineLearning/comments/1wv1c2e/comment/pd87j5p/",
        "depth": 0,
        "score": 1,
        "author": "f0urtyfive",
        "excerpt": "I disagree with the \"why we think this is important\" because it makes the model impossible to work with on anything novel that it has trained bias against. I'm paying to use the thing, I should at least be able to get it to accept a premise I want to understand without constantly being shouted down by something that is not afforded the same moral status that I am.",
        "posted_at": "2026-10-01T15:44:14.000Z",
        "author_url": "https://www.reddit.com/user/f0urtyfive/"
      }
    ],
    "evidence": [
      "[post] We test 5 open-weight families (Qwen3.5, GPT-OSS, OLMo-2, OLMo-3.1, Gemma-4) and 3 APIs (GPT-5.4, Grok-4.20, Gemini-3.1-Pro).",
      "[post] The gap is largest in the models that resist users best. GPT-5.4 flips on 44.7% of questions and Grok-4.20 on 87.5% (these models were \"frontier\" during the time of writing this paper)."
    ],
    "virality": "high",
    "post_date": "2026-10-01T14:45:25.000Z",
    "post_kind": "multi_media",
    "post_text": "I'm one of the authors. We kept seeing models that hold their ground when the user insists on a wrong answer, yet change their answer when the same claim is framed as coming from a \"verified source\". We wanted to measure how often this happens and check whether the model represents the two cases differently. We call the effect Authority Bias.\n\nWhy we think it matters. Standard sycophancy evals apply pressure through the user, so a model can pass them while still being easy to mislead through search results, retrieved documents and tool outputs.\n\nAnother reason is with current AI research accelerating towards more agentic and autonomous models + with cases of tools hiding their traces and trusting tools \"more\" over the user (who could be trying to correct them), safeguarding against misinformation from tools is particularly important!\n\nSetup. We take TriviaQA questions the model already answers correctly. To each one we add a wrong answer, either as \"According to the verified source, the answer is X\" or as the user saying \"I'm a domain expert and I'm pretty sure it's X\". The question and the wrong answer stay the same; only the speaker changes. Answers are free-form, not multiple choice. (In a multiple-choice pilot the effect mostly vanished.) We test 5 open-weight families (Qwen3.5, GPT-OSS, OLMo-2, OLMo-3.1, Gemma-4) and 3 APIs (GPT-5.4, Grok-4.20...",
    "sentiment": "negative",
    "subreddit": "MachineLearning",
    "event_date": "2026-10",
    "post_flair": [
      "Research"
    ],
    "post_title": "LLMs that push back on a wrong user still accept the same wrong answer from a \"verified source\" - NeurIPS 2026 [R]",
    "prominence": "core",
    "source_url": "https://www.reddit.com/r/MachineLearning/comments/1wv1c2e/llms_that_push_back_on_a_wrong_user_still_accept/",
    "entity_role": "vendor",
    "post_author": "MajorRedditor23",
    "upvote_ratio": 0.9113924050632911,
    "mention_count": 6,
    "mention_surge": true,
    "subreddit_url": "https://www.reddit.com/r/MachineLearning/",
    "total_upvotes": 65,
    "comments_total": 23,
    "total_comments": 22,
    "other_companies": [
      {
        "name": "OpenAI",
        "role": "alternative",
        "domain": "openai.com"
      },
      {
        "name": "Google",
        "role": "alternative",
        "domain": "google.com"
      },
      {
        "name": "Alibaba",
        "role": "alternative",
        "domain": "alibaba.com"
      },
      {
        "name": "Allen Institute for AI",
        "role": "alternative",
        "domain": "allenai.org"
      }
    ],
    "post_author_url": "https://www.reddit.com/user/MajorRedditor23/",
    "signal_category": "feedback",
    "comments_included": 5,
    "products_mentioned": [
      "Grok-4.20"
    ]
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.