r/dataengineering
Advice on how to efficiently store and query a large dataset (7tb)
- upvotes
- 35
- comments
- 34
Post
Hi all, hoping to get some advice here as I’ve never dealt with setting up storage and query facility for a dataset as large as this before. I have a large dataset of parquet files, partitioned by date, sat in an S3 bucket. Currently users are utilising Athena to query the dataset via Glue, however as they tend to be querying via an id or via a text search, the results can sometimes take more than 10 minutes just to return one row. I’ve looked into using OpenSearch, which seems to match what I’d need (users will primarily want to do full-text search, the date partitioning is irrelevant) but I’m concerned about the cost, as AI generated estimates are telling me at the very least it would be around 65k a month. Are there any better approaches to this or any routes I should consider? Or is this simply the cost of querying big data?
Extracted from these lines
[comment u/MarchewkowyBog] You might want to look into clickhouse. You didn't really explain what the queries or data is. But because you point to parquet on s3 and athena I presume it's more structured then unsteuctured. Clickhouse has full-text search indecies. Depending on the query it might perform better then elastic or opensearch. If any joins are involved it will most certainly perform better
[comment u/JJGreenerTinejo] Clickhouse