Skip to main content
IBMCustomer feedback

A large international company is migrating a critical accounting process off IBM DataStage due to its complexity, poor documentation, and reliance on a single employee who has since left the team.

What happened

Post: "Datastage to Python/Spark"

Source

RedditSep 17, 2026By u/alone2692

r/dataengineering

Datastage to Python/Spark

upvotes
6
comments
14

Post

Highlighted: the lines this signal was extracted from

Disclaimer: I used AI to translate (it was written in portuguese) most of the text bellow, since I'm not very good in english grammar. I have a degree in software development but little work experience. My team is responsible for a technology-related accounting model at a very large company (not in the US/Europe). Our current process involves around 150 dataframes and 200 jobs, runs on IBM DataStage, and is poorly documented. About 90% of the documentation and business knowledge is in the head of someone who has already moved to another department. Approximately 25% of the department’s monthly budget (around $20 million) is currently not being properly allocated to the correct products. My boss’s boss wants to move this process to a different environment, and we have until the end of this year to understand the current process and build a proof of concept, with the goal of starting the actual review and migration process next year. I managed to replicate one of the jobs using Apache NiFi, but it doesn’t seem like the best approach to me. I’m now looking into using Python/Spark in a Big Data environment. Where would you recommend I start studying? I already have some familiarity with Python and Jupyter. Has anyone here gone through a similar migration from DataStage to Python/Spark?

reddit.com/r/dataengineering/comments/1wj8l0g/datastage_to_pythonsparkRead the full source

Comments on the post

5 of 14 comments
  • “What sources and destinations? Are you doing CDC and transformations too with datastage? To me it seems like Spark might not be the best tool if you’re just doing data movement”

    u/dan_the_lion3 points · Sep 18, 2026View

  • “If you have a $20 million/year budget, why not look at DataBricks or Snowflake for the core stack to use? or are you forced to use full O/S?”

    u/GreyHairedDWGuy1 points · Sep 18, 2026View

  • “You can use Spark if you plan to do heavy transformations on the data.”

    u/jbchand1 points · Sep 18, 2026View

  • “before jumping into spark start by mapping the actual business logic out of the datastage jobs into simple sql or python scripts. 200 jobs with zero documentation is mostly an accounting rules puzzle not a big data problem, especially if the data fits in memory. focus on pyspark dataframes basics and pyspark sql first, and talk to that ex-team member to document the mapping formulas before touchin”

    u/agentUi1 points · Sep 18, 2026View

  • “Are there any experienced people on the team? If you ask me priority nr one, two and three would be business logic. You most likely don't need new architecture for that, you can fix it in datastage. Main question is always what are the benefits of this, what do you plan to achieve using new stack? What will business gain? Well if you just copy the same logic from DS, they gain nothing. Or mayb”

    u/Silly-Swimmer17061 points · Sep 19, 2026View

Extracted by Autobound

From the Signal API record
Signal
Customer feedback

What this signalsUser posts often show product pain before it reaches reviews or churn.

Subreddit
r/dataengineering
Stage
Switching
Event date
Sep 2026

Companies

  • DatabricksAlso named
  • SnowflakeAlso named
  • Apache Software FoundationAlso named

The full record

From the Signal API record

Numbers

Mentions
4

Details

Timing
In progress
Category
Implementation
Virality
Medium
Post kind
Text
Prominence
Core
Company's role
Vendor

Topics and mentions

Topics

  • data migration
  • legacy systems
  • data engineering
  • accounting
  • etl

Flair

  • Help

Products named

  • DataStage

Extraction

Sentiment
Negative
Detected
Sep 17, 2026
signal_type
reddit-company
signal_subtype
customerFeedback

Use this data

Get every Reddit signal for IBM and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at IBM this week?”

  2. Send it to your own tools

    The Signal API returns Reddit signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full reddit-company record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/4cee7f94-392b-5cbb-a079-9afa540f2d99 returns this record as JSON. POST /v1/companies/enrich returns every signal for ibm.com.

{
  "signal_id": "4cee7f94-392b-5cbb-a079-9afa540f2d99",
  "signal_type": "reddit-company",
  "signal_subtype": "customerFeedback",
  "detected_at": "2026-09-17T22:35:50+00:00",
  "company": {
    "name": "IBM",
    "domain": "ibm.com"
  },
  "data": {
    "nsfw": false,
    "stage": "switching",
    "awards": 0,
    "timing": "in_progress",
    "topics": [
      "data migration",
      "etl",
      "legacy systems",
      "data engineering",
      "accounting"
    ],
    "post_id": "1wj8l0g",
    "summary": "A large international company is migrating a critical accounting process off IBM DataStage due to its complexity, poor documentation, and reliance on a single employee who has since left the team.",
    "category": "implementation",
    "comments": [
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wj8l0g/comment/pajjlvm/",
        "depth": 0,
        "score": 3,
        "author": "dan_the_lion",
        "excerpt": "What sources and destinations? Are you doing CDC and transformations too with datastage? To me it seems like Spark might not be the best tool if you’re just doing data movement",
        "posted_at": "2026-09-18T10:06:41.000Z",
        "author_url": "https://www.reddit.com/user/dan_the_lion/"
      },
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wj8l0g/comment/pancdl1/",
        "depth": 0,
        "score": 1,
        "author": "GreyHairedDWGuy",
        "excerpt": "If you have a $20 million/year budget, why not look at DataBricks or Snowflake for the core stack to use? or are you forced to use full O/S?",
        "posted_at": "2026-09-18T20:56:36.000Z",
        "author_url": "https://www.reddit.com/user/GreyHairedDWGuy/"
      },
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wj8l0g/comment/palg5c5/",
        "depth": 0,
        "score": 1,
        "author": "jbchand",
        "excerpt": "You can use Spark if you plan to do heavy transformations on the data.",
        "posted_at": "2026-09-18T16:00:13.000Z",
        "author_url": "https://www.reddit.com/user/jbchand/"
      },
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wj8l0g/comment/pam40k0/",
        "depth": 0,
        "score": 1,
        "author": "agentUi",
        "excerpt": "before jumping into spark start by mapping the actual business logic out of the datastage jobs into simple sql or python scripts. 200 jobs with zero documentation is mostly an accounting rules puzzle not a big data problem, especially if the data fits in memory. focus on pyspark dataframes basics and pyspark sql first, and talk to that ex-team member to document the mapping formulas before touchin",
        "posted_at": "2026-09-18T17:41:42.000Z",
        "author_url": "https://www.reddit.com/user/agentUi/"
      },
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wj8l0g/comment/papw0xo/",
        "depth": 0,
        "score": 1,
        "author": "Silly-Swimmer1706",
        "excerpt": "Are there any experienced people on the team?\n\n If you ask me priority nr one, two and three would be business logic. You most likely don't need new architecture for that, you can fix it in datastage.\n\n Main question is always what are the benefits of this, what do you plan to achieve using new stack? What will business gain? Well if you just copy the same logic from DS, they gain nothing. Or mayb",
        "posted_at": "2026-09-19T06:10:49.000Z",
        "author_url": "https://www.reddit.com/user/Silly-Swimmer1706/"
      },
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wj8l0g/comment/pb81zhu/",
        "depth": 0,
        "score": 1,
        "author": "BillyButteredBooty",
        "excerpt": "You can export datastage jobs to a file. Maybe you can do an analysis with AI?\n\n DataStage is pretty ugly in that it can hide certain logic if you do not know where to look. The export file is a mess but once you understand whats happening it might be easier to migrate.",
        "posted_at": "2026-09-21T19:20:11.000Z",
        "author_url": "https://www.reddit.com/user/BillyButteredBooty/"
      },
      {
        "url": "https://www.reddit.com/r/dataengineering/comments/1wj8l0g/comment/pb8727m/",
        "depth": 0,
        "score": 1,
        "author": "a_bit_of_alright",
        "excerpt": "I would get a migration company to do it on a fixed price basis they’ll be faster and allow you to focus on more important/ valuable activities",
        "posted_at": "2026-09-21T19:41:48.000Z",
        "author_url": "https://www.reddit.com/user/a_bit_of_alright/"
      }
    ],
    "evidence": [
      "[post] My team is responsible for a technology-related accounting model at a very large company (not in the US/Europe). Our current process involves around 150 dataframes and 200 jobs, runs on IBM DataStage, and is poorly documented.",
      "[post] About 90% of the documentation and business knowledge is in the head of someone who has already moved to another department.",
      "[post] My boss’s boss wants to move this process to a different environment, and we have until the end of this year to understand the current process and build a proof of concept, with the goal of starting the actual review and migration process next year."
    ],
    "virality": "medium",
    "post_date": "2026-09-17T22:35:50.000Z",
    "post_kind": "text",
    "post_text": "Disclaimer: I used AI to translate (it was written in portuguese) most of the text bellow, since I'm not very good in english grammar.\n\nI have a degree in software development but little work experience.\n\nMy team is responsible for a technology-related accounting model at a very large company (not in the US/Europe). Our current process involves around 150 dataframes and 200 jobs, runs on IBM DataStage, and is poorly documented. About 90% of the documentation and business knowledge is in the head of someone who has already moved to another department.\n\nApproximately 25% of the department’s monthly budget (around $20 million) is currently not being properly allocated to the correct products.\n\nMy boss’s boss wants to move this process to a different environment, and we have until the end of this year to understand the current process and build a proof of concept, with the goal of starting the actual review and migration process next year.\n\nI managed to replicate one of the jobs using Apache NiFi, but it doesn’t seem like the best approach to me. I’m now looking into using Python/Spark in a Big Data environment.\n\nWhere would you recommend I start studying? I already have some familiarity with Python and Jupyter.\n\nHas anyone here gone through a similar migration from DataStage to Python/Spark?",
    "sentiment": "negative",
    "subreddit": "dataengineering",
    "event_date": "2026-09",
    "post_flair": [
      "Help"
    ],
    "post_title": "Datastage to Python/Spark",
    "prominence": "core",
    "source_url": "https://www.reddit.com/r/dataengineering/comments/1wj8l0g/datastage_to_pythonspark/",
    "entity_role": "vendor",
    "post_author": "alone2692",
    "upvote_ratio": 1,
    "mention_count": 4,
    "mention_surge": true,
    "subreddit_url": "https://www.reddit.com/r/dataengineering/",
    "total_upvotes": 6,
    "comments_total": 15,
    "total_comments": 14,
    "other_companies": [
      {
        "name": "Databricks",
        "role": "alternative"
      },
      {
        "name": "Snowflake",
        "role": "alternative",
        "domain": "snowflake.com"
      },
      {
        "name": "Apache Software Foundation",
        "role": "alternative",
        "domain": "apache.org"
      }
    ],
    "post_author_url": "https://www.reddit.com/user/alone2692/",
    "signal_category": "feedback",
    "comments_included": 7,
    "products_mentioned": [
      "DataStage"
    ]
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.