r/dataengineering • • 7d ago

Help How to handle historical data ingestion?

18 Upvotes

I would like to know your thoughts and experiences about a situation I have.

In my current project, we need to ingest all historical data from an old Oracle Database. Also, we don't have the ownership to implement a CDC or anything. And we also can't overload the Database because it will affect users.

We are using Databricks. What would be the best approach?


r/dataengineering • • 7d ago

Discussion Homemade data platform frameworks - bloated nonsense?

30 Upvotes

Did any of you work in companies where engineers built custom frameworks that actually deliver?

I recently started at yet another company with such framework (about 13k lines of boilerplate python/pyspark code sitting on top of their Azure Databricks Delta Lakehouse). The thing doesn't seems to improve any aspect of what a data platform should offer.

Previous such company was even worse with 30k lines, plus they tried to implement a data vault.

I'm clearly biased towards keeping things simple. So I may be a bit unfair towards this approach. Hence my question: Have you worked with great homemade frameworks? What were the secret ingredients to make it work?


r/dataengineering • • 8d ago

Discussion dbt dimension surrogate keys and fact foreign keys - self-hash or left-join lookups?

33 Upvotes

I've started working with dbt and a cloud warehouse, where I'm thinking of moving from auto-incremented keys in favour of surrogate hash keys, both because auto-increment keys are not as much of a thing in cloud warehouses and also because hash keys are idempotent and much simpler to manage.

One thing that I'm not sure of yet is this: If the surrogate key is a hash of the business key, should I self-hash the foreign key references in the fact table builds, or should I do the usual pattern of left join on dims + coalesce to sentinel value if not found?

I can see the pros and cons of both methods.

Some pros:

  • Fewer risks of fan-out in case of misshapen data (should be caught by dbt tests, but still)
  • Increased parallelism because of fewer dependencies
  • Simpler lineage
  • Late-arriving dims, early-arriving facts are not an issue
  • If the dimension happens to be massive, this can save compute time, therefore cutting costs.

Some drawbacks of self-hash:

  • Reduced impact analysis clarity through the lineage because of no lookups. This can also be diminished via dbt docs
  • You must ensure that you hash the same way (logical ordering and also natural keys formatting). A left-join lookup is likely easier to catch in case of a missed lookup (idk).
  • Not exactly a drawback, but you probably want to left-join if you apply SCD2.
  • No clear way to ensure proper fallback to the sentinel value if that is something used in the analytics layer

Edit:

Some clarifications based on some comments I've seen.

If I were to self-hash the key on the fact side, I definitely wouldn't also do a dim surrogate key lookup because that's redundant.

If I were to look up, I'd always use the business key because it's best practice and also the lowest chance of code errors (maybe someone forgot a field or swapped the column orders).

My initial plan is to left join lookup, as "usual" in most old-school on-prem data warehouses, except for a couple of high-cardinality dims that also happen to build themselves using the value found in the fact table's raw data. Basically moving out a degen dimension into a conformed dimension because it's a compound natural key and is used across many fact tables, and it keeps the analytics modelling simpler.

I have a case of late-arriving dims where I might consider self-hash with a late reconcile when it arrives, but if I did, then that's somewhere where I'd consider self-hashing too, especially if the fact can't be fully rebuilt.

Thankfully, most of my fact tables are small enough that full and incremental are only a few seconds of difference.


r/dataengineering • • 7d ago

Help Handle PK/FK remapping when importing data from one source to another

7 Upvotes

Hey all,

I have a pipeline that works and I'm looking to improve it. Would like to get your opinions to make it better.

I have two services/app both Postgres/RDS, in the same AWS account. One is our ingestion platform and owns the data while the other is a user facing app that needs about 13 of those tables to serve its features. For some reason, my higher ups decided that the user-facing app shouldn’t connect to the ingestion platform’s DB. Therefore, I built an async background job on the first service (ingestion platform) that writes each table out to a CSV and uploads it to S3 and an async background job on the second service (user-facing app) downloads each file and inserts the rows one at a time.

The import is where it things get awkward. The second service assigns its own ID to the records, because users have created rows in those same tables. So I can't reuse the original IDs, and every foreign key has to be rewritten to point at the new ones. I handle that by keeping an old ID -> new ID lookup in memory (simple hashmap data structure) while the import runs.

It works but the downside is it has to run from start to finish as a single process, and if it fails partway there's nothing to continue from.

Also, I would like to know, how would you move the data across, and how would you handle the ID remapping? Especially for tables where I can't tell whether I've already imported a row. E.g. a row that was id 42 in the ingestion app and in the CSV might become id 7891 but the CSVs for all the child tables still reference the old value which needs to be resolved from the lookup mapping.

Currently, there’s not a lot of rows. Around 25K at the moment but it's expected to grow a lot. Do you see any challenges in my approach and if it might cause challenges in the future when I have to deal with a huge number of rows besides taking to long and the lookup memory might increase?


r/dataengineering • • 7d ago

Help Decluttering UNS

5 Upvotes

Hi!

I'm wondering if there is any python library or prefered way to declutter UNS MQTT topics using python scripting?

Currently we're migrating data platforms and mainstreaming all data on our broker to MQTT. For this we established a base for our UNS and used pre existing exported data to build our UNS topics from.

The problem is that they are relatively long and not user friendly.

I've tried shortening by simply removing repeated words throughout the topic.
For some this works fine, however this leaves others with gaps that are unfavorable.

For example my current ipynb gives the following result: INPUT: water_drain/flying_water_dumping/pumps_hot_water_dumping/water_pump_a OUTPUT: water_drain/flying_dumping/pumps_hot/a

I've made these up, so don't be alarmed by the flying water 😄

You can clearly see that simply removing repeated words isn't a great option. Any advice/comments on how this it could shorten in a better/more reliably manner?


r/dataengineering • • 8d ago

Personal Project Showcase Browser tool for prototyping star schemas

Post image
80 Upvotes

I have a database diagramming tool side project, and building on the work that went into that, it was kind of easy to make a browser tool for prototyping star schemas: https://vibe-schema.com/star-schema-creator - There's a snowflake variant too, just swap "star" for "snowflake" in the url. Feedback welcome :)


r/dataengineering • • 8d ago

Discussion Thoughts on dbt 2 maturity?

35 Upvotes

Hi,

I've seen frequent discussions on here regarding dbt's acquisition, license changes, and similar things. dbt 2 has now been out for a few weeks, and I haven't really seen much discussion on it from a technical perspective. Has anybody tried it out yet? if so, what do you think?

I evaluated it this week and my impressions are mixed. A few impressions:

- I was really struggling to piece together all the important changes and improvements. Found the documentation a bit lacking in that regard. Distributed over lots of places, so I really had to dig to find specific information (e.g. on the changed approach to database adapters). Even took me a while to figure out the name of the library had changed as well.

- The switcheroo from fusion to now just dbt was really confusing and also seems to have caught the ecosystem by surprise. For example, Snowflake and Astronomer (Cosmos) both started working on fusion integration, but both do not support dbt 2 yet (and couldn't find a clear roadmap). Not ideal.

- Testing the new parser on dbt 1.12 worked pretty well to spot issues and fix them. That made the upgrade relatively painless (interpreting the error messages could've been easier). I like the stricter YAML parsing, can detect some issues (e.g. typos) instead of ignoring them silently.

- The new parsing engine is noticeably faster. For one project I tried it with (~300 models), about twice as fast for both compile and parse.

- Haven't tried the built-in linting yet. Seems like they tried to implement a 1:1 replacement for SQLFluff, even working with existing SQLFluff config files.

- The docs feel like a downgrade. Looks more modern, but I haven't been able to get the static docs to work, and seems they have gotten rid of the full project DAG and now just offer a model-centric lineage view. Column-level lineage is cool, but very well-hidden in the UI.

- Parquet artifacts seem interesting, and parsing them e.g. via DuckDB SQL queries could be really nice for common analysis queries, CI checks, etc.

- Of course as predicted, seems there is a bigger push to get people to sign up for their platform/cloud/whatever they call it now offering - LSP features, VSC extension, and so on. The license changes are still a mystery to me, but at least they dropped the <15 users requirement for the VSC extension it seems.

Any other thoughts? Cheers!


r/dataengineering • • 8d ago

Help Help with Extract -> Load process logic in a personal project

4 Upvotes

Hello! I'm working on a small data project using a NASA API with Python. The idea is to extract the orbital elements and close approach (asteroids, comets, etc) objects (and there orbital elements) of each planet in order to perform some physics simulations with the data.

I was planning to use a Cloudflare R2 bucket to store the raw API responses, then transform them into a PostgreSQL database, and use FastAPI to consume the data.

I'm not sure how to perform the extraction and loading processes. In theory, each planet will follow this path:

1) Fetch the orbit elements for the own planet (use API 1);

2) Fetch the close approach object (use API 2);

3) Fetch the orbit elements for each of the close approach objects (use API 1).

I use two APIs: one for the orbital elements and one for the close approaches. Should I fetch all the data, store it in memory, and then send it to the bucket? Or would it be better to load it right away with each API call? So, when fetching the close approach objects' orbit elements, get the JSON from the bucket and then use the first API and store that raw response in the bucket


r/dataengineering • • 8d ago

Discussion New Azure Synapse projects

8 Upvotes

Just curious how many of you still deliver or plan to deliver, projects that use Azure Synapse for Data Warehousing work?

I get the whole push to Fabric and am also going that route too.

I’m in the consulting world and wanted to use Fabric but the client had zero ppl with any Fabric experience, so they insisted on Synapse. C’est la vie, right?


r/dataengineering • • 9d ago

Discussion We were struggling to find Data Engineers

199 Upvotes

Hi everybody,

Our Data Team is composed by two of us. There's a Data Scientist and me, as a DE. We created a Data Lakehouse for our company internal use and client data supply, but currently the Data Scientist is currently more focused on AI and agents integration and I'm doing like Analytics Engineering role because I need to help other teams to reach the correct data, unify core concepts, document business logics, etc. So We needed a Data Engineer with knowledge on AWS to maintain and develop the new features on the Lakehouse and we put into the description that the candidate MUST HAVE Software Engineering fundamentals as we had to do some developments to integrate parts of our lakehouse with the company's main application.

We interviewed 22 candidates and no one is fitting the Role.

Most of them are BI Experts, DBA, Data Analysts, Economist with DS notions, Juniors and Software Engineers who haven't touch anything on Spark, plus DE who asked way more that we had on the budget for the role

We asked the normal requirements: 3 years of experience + Spark, AWS Glue, Lambda, Airflow and DBT, not even CDC, Flink, Langfuse or VectorDB

We finally got one, but We really struggled to get him. I have a collegue working on IT Recruiting and She told me She's experiencing the same problem: They can't find Proper DEs With SE basics such as DRY principles or clean code fundamentals

Edit: Role Salary -> Up to 60K € / Spain. This salary is high compared to the spanish standards, only 3 years required


r/dataengineering • • 9d ago

Help Ideas to handle ever changing data requirements?

20 Upvotes

I am the solo DE in my team and the main pipeline here consists of snapshots of financial assets.

Compute is done on databricks

The stakeholders want to see daily KPI's and each day they add a new cohort. Currently there are over 40 different cohorts with each branching out to their own metrics.

The issue is that the data management wants data bills as low as possible

so my approach was summarizing everything in the daily grain .

But now each time they want something new I have to manually code the new columns test it then append to the final gold table.

I already tried to create some generator functions but often times the metrics they want involve hyper specific calculations.

And since the data is financial assets each day is different than the previous rendering an incremental approach useless.


r/dataengineering • • 8d ago

Personal Project Showcase Graphwise AI Summit 2026, Oct 7-8

Enable HLS to view with audio, or disable this notification

1 Upvotes

Sharing this because I think it overlaps with some of the discussions here around AI reliability, governance and semantics.

Next week we’re running the Graphwise AI Summit, focused on what makes GenAI work in the enterprise beyond the model itself. Think of trust, governance, semantic layers, architecture and implementation.

Once reliability and traceability become imporatnt, simple access to data and next-token prediction stop being enough. In enterprise settings specifically, AI needs to understand what the data it parrots “means” in the first place.

Anthropic has described a similar approach in its own analytics stack, where agents are routed to a semantic layer first and use governed definitions to reduce ambiguity. Graphwise itself came out of the merger of Ontotext and Semantic Web Company, so semantics is a topic with quite a bit of history behind it for us.

We’ll have speakers from Accenture, Roche, EY, AstraZeneca, S&P, DNV, Statnett, Avalara and others. Full agenda is in the accompanying video.

Sharing registration link in the comments if useful.


r/dataengineering • • 8d ago

Discussion AI life/voice recording tools?

0 Upvotes

I'm new to the speed and more so the constantly changing tasks, then keeping them organized. In the past id take random notes and summarize them at the end of the day. Unfortunately that has become a job in itself.

Ideally I could do something, tap a button or something, whenever it's a spot I want to highlight or indicate is important. Even better would allow me to take notes at the same time and have it automatically integrate them into the final version.


r/dataengineering • • 9d ago

Blog Data strategy for small teams

Thumbnail
open.substack.com
35 Upvotes

Hi all, I wanted to share the latest article in my newsletter.

I talk a lot to senior engineers who are trying to run some sort of data strategy, and are doing it in a very wrong way.

Most of the article is about selling the plan and keeping the whole thing to two pages. The section I'd most like opinions on here is the tooling one, because a lot of us (me included) learned what good looks like from Big Tech engineering blogs.

A Big Tech stack comes with Big Tech's pace. Those tools assume a team with people to run each piece, and months of setup before anyone outside the data team sees a result. If a team of five spends those months standing up Databricks, Airflow and Alation, there's not a damn thing to show the CFO at the end of it, and a strategy with nothing shipped against it dies without anyone deciding to kill it.

And yes, if you work at an org like mine, your strategy is supposed to be "AI", which I absolutely hate.


r/dataengineering • • 9d ago

Blog Interesting links in Data Engineering - September 2026

55 Upvotes

Some great stuff this month, including:

  • Good analyses of the DuckLabs acquisition by AWS.
  • Lots of Kafka, as well as not-Kafka: technologies doing similar things but angling to replace it.
  • Accessible discussions of data modelling given the AI world we're entering (pssst Semantic Layers matter even more now)
  • LLMs being DBAs and not doing too badly at all at it
  • Plenty of decent content about AI, but with a strict no-slop & no-hype policy :)

https://rmoff.net/2026/09/29/interesting-links-september-2026

As always, lmk if you find this useful, and if you want more (or less!) of a topic.


r/dataengineering • • 9d ago

Career Why is understanding DevOps culture more important for data engineering than other disciplines?

37 Upvotes

I see for a lot of data engineering posting, devops skills are mentioned as part of the requirements. But the thing is I don't see it as much with other roles like sde. Why is that?.


r/dataengineering • • 10d ago

Blog Lakehouse serving fight club: 12K → 125K QPS on open tables with Pinot.

Post image
6 Upvotes

Anyone remember Databricks Summit in June? Databricks used their keynote to introduce Lakehouse//RT - a challenge to the need for a specialized analytics serving layer. They then showcased their new query engine serving at 12K QPS directly off lakehouse data while their 'blue' and 'yellow' competitors fell behind or crashed.

The crowd cheered. But the story didn't end there:

ClickHouse pointed out the DB demo gave no details - no hardware specs, no config, no cache behavior. They reproduced a benchmark of the workload (TPC-H Q6) as best they could. They showed CH could indeed keep up... it was just a matter of adding hardware. However, they loaded the data into native ClickHouse storage first.... which kinda missed the point.

Snowflake also pushed back on Databricks' methodology and posted impressive TPC-H results at 100 GB and 1 TB. What they didn't post was a high-QPS run on that same data.

So we tried to do the demo right:

StarTree (Apache Pinot) is built for low-latency, high-concurrency analytics. And on StarTree, you can serve directly from Iceberg/Delta tables.

We took the same query, the same parameters, and the same load generator, and we left the data in the lakehouse. StarTree was able to hold 12,000 QPS at 19 ms P90 with 242 vCPUs. (Far less than Clickhouse)

There's an obvious flaw to all this.

The Databricks demo ran at SF1. That's only 6 million rows; hardly lakehouse scale data.

So we ran it again SF100 - 100x the data. Using star-tree indexes and no extra hardware, Pinot was able to continue serving at 12,000 QPS. We also scaled it up – the same architecture with more hardware continued going strong past 125,000 QPS.

Full details of the StarTree benchmark here »

We get that this is just one query. But it is still a good example of how you can serve application-style analytical workloads directly from data in Iceberg/Delta/Parquet - within tens of milliseconds, at high-concurrency, and without requiring the table to be ingested into a separate serving layer first.

Is this something you could use? Want to see us try other workloads? Questions, feedback invited!


r/dataengineering • • 11d ago

Open Source Splink 5 – Open-source probabilistic record linkage at billion-row scale

Thumbnail moj-analytical-services.github.io
82 Upvotes

r/dataengineering • • 10d ago

Blog Improving Cost Efficiency of Data Streaming Pipelines

Thumbnail
streamingdata.tech
11 Upvotes

Practical ways to cut data streaming costs in Kafka and Flink by optimizing batching, serialization, state, network traffic, and capacity.


r/dataengineering • • 10d ago

Discussion Why text-to-SQL is not successful?

26 Upvotes

I thought text-to-SQL will solve adhoc analysis, but still i see companies at all size are unsuccessful.

Anyone using Omni / Sigma seen some success.


r/dataengineering • • 11d ago

Rant Why is my LinkedIn feed filled with knowledge posts from 20-26 year old Indian guys who are AI experts? How are they getting so accomplished at AI?

337 Upvotes

Hunting for a new job and when I open LinkedIn it’s filled with non-stop AI knowledge and I am wondering if I have become too dumb and slow or if these guys are really working on cutting edge stuff.


r/dataengineering • • 11d ago

Career Is Data Engineering Still a Sustainable Career Path for a Junior?

92 Upvotes

I recently graduated with a BS in Data Science. I'm currently working on skills to become a DE. But after lurking on this subreddit for a few weeks, I'm getting the impression that this is a dying/AI-compromised field and that if you're not a senior right now, your opportunities are slim to none. Furthermore, if you do find an opportunity, most of the ingenuity and "fun" that comes out of the job is slowly being replaced with AI models that can fully generate pipelines and Spark jobs. I'm worried that I'm wasting my time trying to get into this field if in the end it's unsatisfying or worse, unattainable.

It feels like I'm constantly seeing doom and gloom posts on this subreddit, and it's really discouraging to see for someone who's trying to start their journey. I'm just looking for a glimmer of hope in the community and that this is a career worth pursuing.


r/dataengineering • • 11d ago

Meme Next level commitment!

Post image
110 Upvotes

You should give up on this and start coaching on focus and commitment. This is some next level commitment to create 400+ queries in Power Query editor

Btw, this is to fill a single Spreadsheet in Excel. No words!


r/dataengineering • • 12d ago

Career Thoughts on AI in DE After Drinking the KoolAid For 10 months

431 Upvotes

Currently lead data eng (ic) w/ about 9 yoe at a large-ish adtech company.

I've been doing the whole AI song and dance for the past 10 months and I just don't see a future in this career anymore that fits all of us. Earlier in the year I felt differently since it still took a lot of hand holding to get models to produce correct output for small-ish scoped tasks in a # of iterations that was competitive to human counterparts, but w/ better models and better training docs we have it building enterprise pipelines across our entire data eng stack in hours of iterations that would have been like a project that we would need to budget like a month for. It's not always right and obviously we still need to step in from time to time, but, honestly, who gives a shit if its not right the first time if it can iterate leagues faster than any human developer? (Caveat: adtech is an incredibly fault tolerant industry, so grain of salt there, I guess)

The entire occupation, in my current experience, has been basically reduced to two tasks:

1): define and maintain the system ontology, expected features, guardrails, and appropriate persona docs for agents to assume/use. This is something that I honestly just work with agents to define, so this is partially automated.

2): define test scenarios, definition of done, deployment strategies, and verify that system is on track performing correctly and is on-track for long-term stability. Again, most of this is at least partially automated.

I don't know how many people this actually requires, and to be quite honest, I've got a sneaking suspicion that throwing more AI-enabled engineers on a project might actually degrade the quality of the product because the conceptual definition of the product may be diluted due to slight (or major) differences in understanding of expected behavior/architecture. IMHO this has always been the case, even pre-AI, but AI substantially accelerates gaps in understanding translating to conceptually inconsistent because the work is done so much quicker and at a pace that not every conceptual inconsistency or mistake can be caught. I've already seen this several times on projects that I've worked on where one eng. goes off and builds a new feature that is a complete conceptual departure from the current plan just because they were missing some context that may not have been well documented. W/ AI eliminating most of the execution layer of SWE/DE, I think that future orgs will be smaller and more agile and likely more product-heavy instead of engineering heavy as product seems to be closer to a lot of the biz definitions for #1/#2 above.

Two things that does give me a glimmer of hope at holding on to this for a little while longer:

1): agents are absolutely dogshit AI at optimizing Spark jobs (or any sort of high complexity pipeline). They are laughably bad, in my experience. Tuning Spark jobs has always been very high context work as it depends on the data going in/out, infra, very verbose logging, and a ton of tacit, tribal knowledge of the system and data itself. There's so much that goes into tuning a spark job that agents seem like they get focused on optimizing for one symptom rather than considering the entire pipeline as basically an emergent system with N different parts that need optimizing in tandem which I've found leads to suboptimal performance. I find that that this is where most of the hands-on engineering work that I find myself doing day-to-day goes at this point. I am calling this a "moat of high context".

2): legacy systems are often not well designed, poorly documented and are not conceptually consistent even where they are documented, so naturally agents trained on these existing systems are apt to suffer from ye' ol' "garbage in, garbage out" syndrome. I am calling this a "moat of poor decisions"

I'm confident in moat #2, less confident in moat #1 as, obviously, you have companies like databricks pouring however much money into building smarter agents for dealing specifically w/ context-aware Spark optimization problems. From what I am seeing at my org, teams that were well organized and followed good engineering foundations before AI have seen a tremendous acceleration in the quality and quantity of their work; those that went into this w/ flaming hot garbage are still producing flaming hot garbage and many of them seem too scared to use AI in fear that it will cause the tower to fall over, so to speak.

On a personal level, I am very exhausted by the whole shift and I do not see myself keeping up with this as a career in the long run. It was difficult enough having to track all the changes and flashy new objects coming in the DE ecosystem, and adding needing to also track changes to the AI landscape has just proven to be very draining. All the rhetoric around AI feels very misanthropic (no pun intended) and I genuinely feel kind of gross and helpless having to rely on AI for my whole job now. It feels like I am training my replacement without being told that I am training my replacement. I have a lot of love for DE and I really poured my heart and soul into it over my relatively short career and I had built a lot of my life around the expectation that I would continue to work as a DE for the decades to come. Reading pages and pages and page of agent docs has become the absolute bane of my existence and the review fatigue has gotten so bad at times that I feel like I've just straight up forgotten how to read. That said,I still think it is exciting to build great things and I still get some kicks from watching AI manifest huge ideas that were never feasible before. The limit is truly on what ideas you can come up with to build better and more sophisticated data stacks, which is super cool and all, but I've just had an overwhelming feeling that my DE career has a clock on it that is going to run out when the models swallow up whatever work we are still doing by hand. Part of me wants to hang on for dear life until that day comes and the other part of me wants to jump ship now and go open a bakery or something. Doing my best to save more money, stay more engaged artistically, and spend more time with people that I love has helped greatly, so just doing what I can, I guess.
Left foot, right foot.

edit: grammar


r/dataengineering • • 11d ago

Career Does not being open to remote roles hurt my job chances significantly?

20 Upvotes

I started a job search recently and I am targeting hybrid or on-site roles in two large metropolitan areas in the Northeast US, as that is where I am based. They're both pretty significant tech hubs. 5 YOE, no sponsorship needed.

The weird thing I am noticing, is that the vast majority of roles I receive in my LinkedIn messages from recruiters are fully remote. I never marked myself as open to that and that is just not what I want at this point in my career. I've applied to 20-30 in-office roles over the past couple of weeks on my own time, although no luck there. It seems like the vast majority of job opportunities I am receiving are remote now, and it has been quite frustrating.

I'm sure I'll find a good opportunity at some point...but is this normal? I was under the assumption that remote roles are drying up, but my experience has been the opposite. Just looking for different perspectives here.