Skip to main content
ClickHouseCustomer feedback

Multiple commenters recommend ClickHouse as a high-performance alternative to OpenSearch or Elasticsearch for querying a 7TB structured dataset, highlighting its full-text search indices and...

What happened

Multiple commenters recommend ClickHouse as a high-performance alternative to OpenSearch or Elasticsearch for querying a 7TB structured dataset, highlighting its full-text search indices and superior performance on joins.

Source

RedditSep 22, 2026By u/Expanding_Dong_

r/dataengineering

Advice on how to efficiently store and query a large dataset (7tb)

upvotes
35
comments
34

Post

Hi all, hoping to get some advice here as I’ve never dealt with setting up storage and query facility for a dataset as large as this before. I have a large dataset of parquet files, partitioned by date, sat in an S3 bucket. Currently users are utilising Athena to query the dataset via Glue, however as they tend to be querying via an id or via a text search, the results can sometimes take more than 10 minutes just to return one row. I’ve looked into using OpenSearch, which seems to match what I’d need (users will primarily want to do full-text search, the date partitioning is irrelevant) but I’m concerned about the cost, as AI generated estimates are telling me at the very least it would be around 65k a month. Are there any better approaches to this or any routes I should consider? Or is this simply the cost of querying big data?

Extracted from these lines

  • [comment u/MarchewkowyBog] You might want to look into clickhouse. You didn't really explain what the queries or data is. But because you point to parquet on s3 and athena I presume it's more structured then unsteuctured. Clickhouse has full-text search indecies. Depending on the query it might perform better then elastic or opensearch. If any joins are involved it will most certainly perform better

  • [comment u/JJGreenerTinejo] Clickhouse

reddit.com/r/dataengineering/comments/1wnmevk/advice_on_how_to_effici...Read the full source

Comments on the post

5 of 34 comments
  • “use icerberg open table format.”

    u/BlackHole_Rider37 points · Sep 23, 2026View

  • “usually the tool is not the answer, first think if you can better partition the data, is it possible to create it in such a way that the query engine can easily filter on results and thus query less? Are all results always necessary or can you further archive old data, can you filter on id ranges, etc etc.”

    u/Slampamper14 points · Sep 23, 2026View

  • “It's pointless to propose solutions if you dont give access pattern and SLA (ie it should come back in below 10ms). ElasticSearch is a particularly bad advice if you just need to query by an id, this would be crazy wastefull.”

    u/fdqntn14 points · Sep 23, 2026View

  • “Partition your data based on the query patterns”

    u/notmarc113 points · Sep 23, 2026View

  • “You might want to look into clickhouse. You didn't really explain what the queries or data is. But because you point to parquet on s3 and athena I presume it's more structured then unsteuctured. Clickhouse has full-text search indecies. Depending on the query it might perform better then elastic or opensearch. If any joins are involved it will most certainly perform better”

    u/MarchewkowyBog11 points · Sep 23, 2026View

Extracted by Autobound

From the Signal API record
Signal
Customer feedback

What this signalsUser posts often show product pain before it reaches reviews or churn.

Subreddit
r/dataengineering

Companies

  • OpenSearchAlso named
  • ElasticsearchAlso named
  • Amazon Web ServicesAlso named

The full record

From the Signal API record

Numbers

Mentions
1

Details

Category
Features
Virality
High
Post kind
Text
Prominence
Aside
Company's role
Vendor

Topics and mentions

Topics

  • data warehousing
  • query performance
  • database
  • full-text search

Flair

  • Help

Extraction

Sentiment
Positive
Detected
Sep 22, 2026
signal_type
reddit-company
signal_subtype
customerFeedback

Use this data

Get every Reddit signal for ClickHouse and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at ClickHouse this week?”

  2. Send it to your own tools

    The Signal API returns Reddit signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full reddit-company record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/28580f4d-9db1-5575-a6e3-6aaf93ebd35c returns this record as JSON. POST /v1/companies/enrich returns every signal for clickhouse.com.

{
  "signal_id": "28580f4d-9db1-5575-a6e3-6aaf93ebd35c",
  "signal_type": "reddit-company",
  "signal_subtype": "customerFeedback",
  "detected_at": "2026-09-22T21:20:41+00:00",
  "company": {
    "name": "ClickHouse",
    "domain": "clickhouse.com"
  },
  "data": {
    "nsfw": false,
    "stage": "none",
    "awards": 0,
    "topics": [
      "data warehousing",
      "query performance",
      "database",
      "full-text search"
    ],
    "post_id": "1wnmevk",
    "summary": "Multiple commenters recommend ClickHouse as a high-performance alternative to OpenSearch or Elasticsearch for querying a 7TB structured dataset, highlighting its full-text search indices and superior performance on joins.",
    "category": "features",
    "comments": [
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wnmevk/comment/pbj4uzm/",
        "depth": 0,
        "score": 37,
        "author": "BlackHole_Rider",
        "excerpt": "use icerberg open table format.",
        "posted_at": "2026-09-23T09:08:02.000Z",
        "author_url": "https://www.reddit.com/user/BlackHole_Rider/"
      },
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wnmevk/comment/pbji9b3/",
        "depth": 0,
        "score": 14,
        "author": "Slampamper",
        "excerpt": "usually the tool is not the answer, first think if you can better partition the data, is it possible to create it in such a way that the query engine can easily filter on results and thus query less?\nAre all results always necessary or can you further archive old data, can you filter on id ranges, etc etc.",
        "posted_at": "2026-09-23T10:54:42.000Z",
        "author_url": "https://www.reddit.com/user/Slampamper/"
      },
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wnmevk/comment/pbj8vc6/",
        "depth": 0,
        "score": 14,
        "author": "fdqntn",
        "excerpt": "It's pointless to propose solutions if you dont give access pattern and SLA (ie it should come back in below 10ms).\n\n ElasticSearch is a particularly bad advice if you just need to query by an id, this would be crazy wastefull.",
        "posted_at": "2026-09-23T09:42:34.000Z",
        "author_url": "https://www.reddit.com/user/fdqntn/"
      },
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wnmevk/comment/pbjq3z4/",
        "depth": 0,
        "score": 13,
        "author": "notmarc1",
        "excerpt": "Partition your data based on the query patterns",
        "posted_at": "2026-09-23T11:45:43.000Z",
        "author_url": "https://www.reddit.com/user/notmarc1/"
      },
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wnmevk/comment/pbiyi01/",
        "depth": 0,
        "score": 11,
        "author": "MarchewkowyBog",
        "excerpt": "You might want to look into clickhouse. You didn't really explain what the queries or data is. But because you point to parquet on s3 and athena I presume it's more structured then unsteuctured. Clickhouse has full-text search indecies. Depending on the query it might perform better then elastic or opensearch. If any joins are involved it will most certainly perform better",
        "posted_at": "2026-09-23T08:11:23.000Z",
        "author_url": "https://www.reddit.com/user/MarchewkowyBog/"
      },
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wnmevk/comment/pbjtv3m/",
        "depth": 0,
        "score": 5,
        "author": "ludflu",
        "excerpt": "you should probably think about access patterns and typical queries before jumping straight to the tech. What are the kinds of questions do people want to answer using this data?\n\n Use those questions to help understand how you should structure the data, then use that structure to help decide what tech would be a good fit.",
        "posted_at": "2026-09-23T12:07:41.000Z",
        "author_url": "https://www.reddit.com/user/ludflu/"
      },
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wnmevk/comment/pbiybsr/",
        "depth": 0,
        "score": 4,
        "author": "JJGreenerTinejo",
        "excerpt": "Clickhouse",
        "posted_at": "2026-09-23T08:09:52.000Z",
        "author_url": "https://www.reddit.com/user/JJGreenerTinejo/"
      },
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wnmevk/comment/pbjgqfj/",
        "depth": 0,
        "score": 1,
        "author": "Difficult-Bag1550",
        "excerpt": "if users query by id, you could add bucketing to this column",
        "posted_at": "2026-09-23T10:44:01.000Z",
        "author_url": "https://www.reddit.com/user/Difficult-Bag1550/"
      }
    ],
    "evidence": [
      "[comment u/MarchewkowyBog] You might want to look into clickhouse. You didn't really explain what the queries or data is. But because you point to parquet on s3 and athena I presume it's more structured then unsteuctured. Clickhouse has full-text search indecies. Depending on the query it might perform better then elastic or opensearch. If any joins are involved it will most certainly perform better",
      "[comment u/JJGreenerTinejo] Clickhouse"
    ],
    "virality": "high",
    "post_date": "2026-09-22T21:20:41.000Z",
    "post_kind": "text",
    "post_text": "Hi all, hoping to get some advice here as I’ve never dealt with setting up storage and query facility for a dataset as large as this before.\n\nI have a large dataset of parquet files, partitioned by date, sat in an S3 bucket. Currently users are utilising Athena to query the dataset via Glue, however as they tend to be querying via an id or via a text search, the results can sometimes take more than 10 minutes just to return one row.\n\nI’ve looked into using OpenSearch, which seems to match what I’d need (users will primarily want to do full-text search, the date partitioning is irrelevant) but I’m concerned about the cost, as AI generated estimates are telling me at the very least it would be around 65k a month.\n\nAre there any better approaches to this or any routes I should consider? Or is this simply the cost of querying big data?",
    "sentiment": "positive",
    "subreddit": "dataengineering",
    "post_flair": [
      "Help"
    ],
    "post_title": "Advice on how to efficiently store and query a large dataset (7tb)",
    "prominence": "aside",
    "source_url": "https://www.reddit.com/r/dataengineering/comments/1wnmevk/advice_on_how_to_efficiently_store_and_query_a/",
    "entity_role": "vendor",
    "post_author": "Expanding_Dong_",
    "upvote_ratio": 0.926829268292683,
    "mention_count": 1,
    "mention_surge": false,
    "subreddit_url": "https://www.reddit.com/r/dataengineering/",
    "total_upvotes": 35,
    "comments_total": 34,
    "total_comments": 34,
    "other_companies": [
      {
        "name": "OpenSearch",
        "role": "competitor",
        "domain": "opensearch.org"
      },
      {
        "name": "Elasticsearch",
        "role": "competitor",
        "domain": "elastic.co"
      },
      {
        "name": "Amazon Web Services",
        "role": "alternative",
        "domain": "amazon.com"
      }
    ],
    "post_author_url": "https://www.reddit.com/user/Expanding_Dong_/",
    "signal_category": "feedback",
    "comments_included": 16
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.