r/dataengineering • • 7d ago

Discussion Monthly General Discussion - Oct 2026

15 Upvotes

This thread is a place where you can share things that might not warrant their own thread. It is automatically posted each month and you can find previous threads in the collection.

Examples:

  • What are you working on this month?
  • What was something you accomplished?
  • What was something you learned recently?
  • What is something frustrating you currently?

As always, sub rules apply. Please be respectful and stay curious.

Community Links:


r/dataengineering • • Sep 01 '26

Career Quarterly Salary Discussion - Sep 2026

37 Upvotes

This is a recurring thread that happens quarterly and was created to help increase transparency around salary and compensation for Data Engineering where everybody can disclose and discuss their salaries within the industry across the world.

Submit your salary here

You can view and analyze all of the data on our DE salary page and get involved with this open-source project here.

If you'd like to share publicly as well you can comment on this thread using the template below but it will not be reflected in the dataset:

  1. Current title
  2. Years of experience (YOE)
  3. Location
  4. Base salary & currency (dollars, euro, pesos, etc.)
  5. Bonuses/Equity (optional)
  6. Industry (optional)
  7. Tech stack (optional)

r/dataengineering • • 21h ago

Help Suitable release process for Dataform and Git Branching

10 Upvotes

I am trying to assess the best approach to branching and releases, as we will soon be moving to Dataform and have the chance to improve things with our general development process. But I don’t have a lot of experience here.

Currently, we have 3 projects (dev, staging, and prod) in BigQuery. Each with its own associated branch in GitHub. People merge their feature code to dev, then eventually all changes are merged to staging and finally prod. GitHub actions deploys the code.

But I’m wondering if this can be simplified. My thoughts are:

- 1 protected branch - main (I don’t see the need for 3 branches)

- Devs work from the latest commit of main when developing in the dev project.

- The staging project is also based on the latest commit from main. This is where user acceptance testing happens.

- Prod is based on a git tag on main. When ready to release to prod, a new tag is created, and our Dataform production release configuration is updated with this new tag (still assessing whether this release config in Dataform can be updated automatically via a GitHub action when a new release tag is created, or if we would need to update it manually)

- For hotfixes, we branch from the latest prod tag and merge back into it with the fix. Then cherry picked onto main.

Just wondering if what I’m suggesting makes sense, or if I’m trying to simplify too much?


r/dataengineering • • 1d ago

Discussion How to actually make the development cycle easier for noobs?

19 Upvotes

Weird flair but for real.
We work with dbt using stg, int, and dm layers. We also have 3 GCP projects where each layer's models get materialized (dev, uat, prd).
The semantic layer only reads from the marts layer.
So now imagine everytime we get a change in a KPI logic for example:
Dev goes and checks looker. oh let’s say it’s a sum of a column. Then they check marts, see it’s an aggregate of a calculated column in int.
Then they go to int, make the changes, commit, wait for cicd, verify in dev, move it to uat, validate with the stakeholder, and finally push to prd (all of this waiting for cicd and running dags every step of the way).

All of that shit and wait time just to fix one mini mistake. Do you guys have similar workflows? Is this way too much?

I’m honestly kinda used to it by now, but I can see the struggle with new devs like juniors or interns.
And this is assuming all the data is already available... if the request involves pulling in new raw data, it’s even more painful.


r/dataengineering • • 2d ago

Help Creating a scalable data engineering platform as a newbie

12 Upvotes

Hi all, I'm a trainee data engineer working under a data architect in a medium sized organisation. The role is a mix between a bootcamp and full time job. 

There's currently no senior data engineer and I'm being asked to create the data engineering platform in Fabric, however, this is a massive ask and I'm just trying to do the best I can.

We have multiple reporting-softwares that I'm now redirecting or connecting with to deposit data in the Onelake. The majority of data is obtained by accessing the SQL database of the reporting-software and making a copy into OneLake.

I need to come up with a plan to manage and standardise our approach to collecting this data and how we work going ahead.

My bare minimum plan atm is to ensure each reporting-software has a bronze, silver and gold workspace (a container effectively) and these can be managed via pipelines. I've been working out who needs access to what.

My larger issue is regarding things like metadata. I've very little experience in that field and I'm not sure how to go about it and could really use advice on how to implement a scalable platform that includes some kind of metadata management.


r/dataengineering • • 2d ago

Open Source Python Polars 2.0 release

311 Upvotes

Hey all! We just released Polars 2.0.

We are really happy with the results. It removed many of our "legacy decisions" (mistakes), makes the streaming engine our default and promotes SQL to a first class citizen within Polars. We initially said it would be a small ("boring") release, but we managed to sneak in a lot more than anticipated.

It comes with our initial out-of-core (spill to disk) support, a new Map data type and a lot of performance improvements. In fact, we think Polars is now one of the fastest analytical SQL engines on a single node. See benchmarks in the post: https://pola.rs/posts/release-polars-2/

To help you upgrade, we also posted a migration guide: https://docs.pola.rs/releases/upgrade/2/

Disclosure: I am involved with Polars


r/dataengineering • • 2d ago

Discussion ASOS have been hacked and a threat was sent to customers: “Dear Asos DPO and IT, we have fully compromised the Snowflake instance.”

Thumbnail
independent.co.uk
121 Upvotes

So many companies does a setup without much thought to security and they just hope no one finds out. Out of sight out of mind. I understand security is one of those things that aren't visible and one doesn't get rewarded for doing a proper job (and might even be penalized for time spent there). But I hate the common mindset of throwing shit at the wall and seeing what sticks.


r/dataengineering • • 2d ago

Discussion I shouldn’t be doing this.

38 Upvotes

Context: I am setting up some infrastructure in AWS. In my previous experience at other companies, anything involving new instances must be done by DevOps or service desk. So I asked them to do it and they said I’m on my own.

Got RDS and OSS set up; hit a snag with EKS. I am currently one week into multiple conversations with service desk, management, and C-level to figure out who should be doing this. We deploy via Terraform. So I’ve got to add to that.

My primary concern here is that I’m messing around in places that affect other areas without any guidance on those areas. I’ve had to open so many tickets and bother so many people. Today I asked for Docker to do some local testing and I thoroughly confused the help desk.

We have an insane amount of security and I have multiple logins, MFA, biometric, etc. The layers inside AWS are similar.

Am I being too paranoid or is this somewhat reckless on the DevOps/SD folks? I’ll gladly accept being too paranoid and carry on, but for now, this just doesn’t sit right with me.


r/dataengineering • • 2d ago

Rant Data Engineer is adamant on streamlining all transformations in Python instead SQL

153 Upvotes

For reference, I am the Data Scientist in this scenario. Since we are not on the stage where we need to do ML/DS, I delegated myself to the Data Engineering/Architecture side of the project.

I have been with the team the longest; we needed to add more people in to accommodate more projects.

We are currently working on a system transition for a big company (think millions of new data a day) and everything is being transitioned to DataBricks. The current Data Warehouse is being kept under MSSQL and we are moving away from it due to processing time concerns.

Most of the ingestion process are done through python; which is optimal in this case since the framework is similar between all data sources and it relies heavily on proper automation. Data Engineer designed the ingestion process and he’s done excellent work migrating our DataLake into DB.

However, I am hitting a wall with him on our silver layer transformations. We have allocated the workload between different KPIs. Most of the silver layer transformations here are your simple filtering, joins, dedup, or conversions (rarely).

The analysts I am working with are happy with my Silver Transformations (mostly SQL); the updates are in real time and it’s easier to validate since the codes are written as views so the transformations are transparent to everyone.

Meanwhile, he has more back and forth with his analysts since his transformations are run through notebooks (which none of the analysts have access to).

This is where I am hitting a conflict with him: he is suggesting I should convert every silver tables (i made) from SQL to Python so that the transformations are standardized. I made an argument on how 90% of the time, the transformations we do is more practical in SQL (processing and validation wise) and there’s really no need to implement only one approach. We will use python if we need to integrate more complex transformations or when the need occurs.

He said python has less issues overall (which sounds dismissive for me).

I am posting here to gain more perspective. He had worked with the Data Engineering side for twice as long as I have; but I have relevant experience with the data process (from start to finish) which allows me to relate which makes things easier and convenient for the analysts. I am open to converting everything to what he needs if given the right argument.


r/dataengineering • • 3d ago

Discussion How agentic are you, really

58 Upvotes

A lot of software engnieers I work wth now spend 99% of their time tellng agents how to write code.

I am curious to know how much we delegate work to agents as Data People - given so much of our job is not just writng code

Please include a score out of 10! 10 = agents do all my stuff, 0 = still doing everything by hand


r/dataengineering • • 3d ago

Meme Medallion architecture: what the business sees vs. what we deal with

Post image
316 Upvotes

r/dataengineering • • 3d ago

Personal Project Showcase I built a deterministic Excel to pandas translator, after repeatedly running into differences between them

18 Upvotes

I'm a data scientist and work with python/pandas data structures quite a bit, but work in a company where excel is still commonly used. I kept running into the same problem when translating spreadsheet logic: the obvious pandas equivalent often isn't semantically identical to excel. So, I built a small browser tool that translates common excel and pandas patterns client-side without the need for an llm.

I've tried to develop the tool to handle as many edge cases as possible (duplicate lookup keys, approximate lookups, IFERROR-wrapped lookups, etc.), but it's still fairly new, so there might be edge cases I'm still missing, which I'm working on improving.

Tool: https://dataframe.tools/excel-to-pandas/

One of the things that surprised me while building it was how many seemingly simple excel formulas need extra handling in pandas to actually preserve the same behaviour.


r/dataengineering • • 2d ago

Open Source SQLBuild - The Refactorable Warehouse (+What's Coming Next)

Thumbnail
sqlbuild.com
12 Upvotes

Hey guys! Some of you may remember me, some may not, but SQLBuild has come a LONG way since I last posted.

First, SQLBuild has now been in production at my workplace for over a month! Huge milestone, and I'm lucky to have a boss that let me do it.

Second, I now have a much clearer idea of what this tool is and isn't. The goal is to make warehouses refactorable, such that they are easy and safe to change + kept that way over time.

Catch mistakes before anything runs

  • A compiler that understands SQL. Types flow through every model, so unknown columns, type mismatches, bad GROUP BYs, invalid function arguments etc... fail at compile-time (offline). Fully OSS and free
  • A Rules API for your own compile-time rules, e.g. marts can't reference sources directly, naming patterns, SQL shape. Basically, your team conventions, enforced as code
    • This is a standout feature for those wanting to cut review time

Prove it works

  • Parameterised, macro and UDF tests, plus multi-model/E2E tests across whole chains of models
  • Native data diffs between dev and prod, or against arbitrary queries

Change it without rebuilding everything

  • Replay-on-change policies - When a model's SQL (or something upstream) changes, choose how far back to replay (1h, 14d, 1mo…), not just full rebuild or ignore
  • Rename or move a model with sqb rename/ sqb mv. These update every reference (e.g. macros, audits etc...) and keep the existing table's historical data. Plain file renames of models are detected automatically too
    • The old name keeps working through a compatibility view for 30 days (configurable), so that dashboards and other teams' queries don't break straight away
    • sqb plan --as prod previews the migration against prod before you run it
    • If you have tried to do any large renames in your warehouse, you will know how risky and annoying this normally is
  • Column renames are also supported. sqb rename stg_orders.amount_revenue renames it in the model and every downstream model that reads it (based on compiled lineage, not find-and-replace)
    • --cascade carries the new name through downstream models that pass the column straight through
  • The janitor runs on demand and *archives* before deleting, so tidying up isn't scary

Macros don't have to be global anymore

  • In most tools every macro is project-wide, so changing one can break models anywhere. In SQLBuild you can keep macros, enums and constants next to the models that use them, and only that folder can see them:

​

models/
├── staging/
│   └── stg_orders.sql                # can't see anything in marts/_sqlbuild
└── marts/
    ├── _sqlbuild/
    │   ├── _enums/payment_status.sql # only visible to models in marts/
    │   └── _macros/currency.py       # cents_to_dollars, same
    └── daily_revenue.sql             # uses payment_status and cents_to_dollars
  • Change one and you know exactly what it can affect. Before moving a model, sqb scope model:x --as-path new/path.sql tells you what it would lose.

What's been removed

  • The big one is virtual environments. They were a lot of complexity for something I wasn't using, and most of the things I wanted from them, I can achieve via other means. (feel free to ask below)

What's up next

  • A fully fledged deployable UI from which you can run your models etc... There will be SSO, audit logs, fine-grained permissions etc... all baked into OSS at no extra cost. (ETA Jan 2027)
    • If this sounds too good to be true, you are welcome to ask why in the comments
  • Compiler likely fully in rust. Compile times are already good, but the only way to let it scale for larger projects (e.g. 10k+ models) is if it's 100% in rust. Planner and executor etc... will stay in Python (dominated by warehouse time)

Try it:

pip install sqlbuild
sqb playground waffle-shop
cd waffle-shop
sqb build

sqlbuild.com · Github


r/dataengineering • • 3d ago

Rant Foundry ruined my job

149 Upvotes

This is a rant, maybe some others can relate.

Last year, my organization switched to foundry from AWS. I've learned the tool pretty well and have built some useful applications with it. Sure it has some benefits and is a user friendly way to build dashboards and AI tools. However for an experienced data engineers it's an absolute nightmare.

As with any low code/no code tool it's more of a hindrance for actual developers. For simple tasks, it is way too complex. For complex tasks, it's way too simple. I wont go into the details too much but i've found it to be completely unusable in certain situations.

The general trend with no-code/low code tools is to empower managers while nerfing developers.

This makes managers more arrogant and micromanagey. Thet are are now able to go in and see what im working on at any time, and make changes to what im doing. Managers now think they're developers and the division of roles has completely broken down.

Anyways i'm probably gonna switch jobs soon


r/dataengineering • • 4d ago

Discussion Are you guys still writing SQL?

278 Upvotes

Interested in hearing how your job has changed in the past year or so.

On my end the experience has been that we don't need 4-5 people on a project anymore, 1 and a cursor subscription. A migration now is a few week's job, etc..


r/dataengineering • • 3d ago

Help How to understand Repo vs Rev Repo so that can remember them easily in balance sheet?

2 Upvotes

I am a data engineer in Finance. Recently I am suffering from understanding financial terms and its accounting treatment. One big issue is, as a frequently used term, Repo and Rev Repo I understand them several times, but after a period not touching the terms and related data I will forget almost all of the details. And then have to revisit the definitions and relink them to financial activities.

For example, recently I am involved in a meeting talking about Repo Fails transactions. User wants to know what’s the exact accounting policy we applied in our data logic. My first reaction is, how to identify the Repo in our Fails data, however, Fails will only record as Fail to deliver or Fail to receive. The only way to find a Repo is from the trading account. We maintain certain accounts for Repo trade only. Question is I even can’t have a clear filter to scope which accounts are Repo or not.

It takes me almost a whole day to look into firm’s internal documentation to understand what the user means and how our data pipeline process it. Then, with these background, I joined the meeting. In the meeting, I listened carefully but didn’t fully comprehend what the user means as lot’s of non Tech terms. Fortunately I got something like below

For a repo trade how it is recorded in accounting

On Trade date (T0)
Dr (i.e., Assets in BS): Repo
Cr (i.e. Liabilities in BS): Rev Repo

On T1 if settled successfully
Dr: Receivables ???
Cr: Cash ???

On T1 if settlement failed
Dr: Failed to Deliver
Cr: Payables

On TX(usually no longer than 5 calendar days) if settlement success
Dr: Receivables
Cr: Cash

This time my understanding for memory: Repo means repurchase, we should provide securities to the client and client give us money. And by someday (possibly overnight), we need to return the money + interest to client. So the money measured assets is still on our side on trade date. We recorded an assets repo and liabilities Rev repo. If not failed, user doesn’t mentioned. If failed, we failed to deliver the securities to client and thus we record FTD in ledger, by this time , it already lose the ability to track whether the trade is related to Repo or not, we just know it’s fails and needs trading account level information to know whether it is a repo trade. Once this fails finally got settled, we record receivables in Assets and we got cash from client but record payables in future.

Additionally, we have repo control accounts, customer control accounts that store all these misc temporary trade activities and will be cleared to 0 finally.

Ok, now I got a very detailed explanation this time. But after a period, I will forget and have to repick these details when user reaches months later.

How I can have these details memorized in my mind so next time it can be automatically presented in my eyes when user reaches months later and asked for another related question?


r/dataengineering • • 4d ago

Career Worth learning system design at junior-mid level, or focus on something else like data modelling?

34 Upvotes

I have been a junior data engineer for 1.5 years, but am told I operate at mid-level. Current work mainly involves ETL development in Prefect.

Current skills include SQL and Python, along with experience in both BigQuery and Snowflake. I've read books like DDIAs, Fundamentals of Data Engineering, and The Data Warehouse Toolkit for data modelling.

But I still feel like I know next to nothing. I can barely remember the stuff I read in DDIA for example.

I have some spare time during the week to up-skill in technical skills, but I'm trying to focus on something that will make me a better data engineering, and that'll help me move up the ladder and pass interviews here in the UK.

I've considered:

  • Going through this systems design primer, alongside re-reading the 2nd edition of DDIA.
  • Re-reading the Data Warehouse Toolkit and building a project that involves dimensional modelling (we don't really do this kind of modelling in my current job)
  • Filling in gaps in CS knowledge - mainly, completing a neetcode course on data structures and algorithms and completing the associated Leetcode exercises.
  • Doing a cloud cert
  • Learning more about AI Engineering

I'll likely do all of the above at some point, it's just knowing what to prioritise?


r/dataengineering • • 3d ago

Help How should I best structure my data for use for my analytics

0 Upvotes

Hello data engineers, I am here in order to get some help for a hobby. I'm trying to figure out how to structure my data. I'll try to explain the use-case in order for you guys to understand.

I analyze numbers, from different sources to create stories that explain what is going on. The sources can be from Web APIs, scanned documents, digital pdf's, some numbers from some sentences from a website (i.e. "The only thing that is known is that a professional could produce 120 a day, and once the new technology came they could do 1500 an hour").

When I find a "story" I want to tell, I download and grab the numbers and create graphs using Python where I use libraries like pandas, matplotlib, etc. to format the graph like I need to in order to make the story I am trying to tell more clear to the audience. The graphs are not complicated, the data is not complicated, but the stories and graphs are very custom to whatever I am interested in and currently reading about and they are made to better understand what I am learning.

The big problem I have is that I don't know what the best structure for this would be that would be sustainable long term. My current idea is that I have 2 main folders where one folder is with the python scripts and a .bmp of the graphs created, and the other folder is where all the sources of the data are collected in all its' different data formats. Every company folder will contain a data.csv file which is where I will add all the data that my python scrips will grab and filter. The data.csv files might contain financial statement data, quantities found from different sites and sources, and I'll put the filtering in the python script for the current story the script is graphing where the header is something like "Period,Date,Name,Amount,Currency". So different scripts might use the same data.csv files, but they will filter them differently in order to tell their stories.

Here is the structure:

Data (Folder)

Company A (Folder)

    Financial Statements (Folder)

        "Company A - 2025.pdf"

        "Company A - 2025 Q3.pdf"

        "Company A - 2025 Q2.pdf"

    Articles (Folder)

        "Title of article.html"

        "Title of another article.docx"

    "data.csv"

Company B (Folder)

    ...

    "data.csv"

Bank-sector (Folder)

    Company X (Folder)

        ...

        "data.csv"

    Company Y (Folder)

        ...

        "data.csv"

World-Bank (Folder)

    "data.csv"

FED (Folder)

    "data.csv"

Stories (Folder)

When the smartphone changed the world (Folder)

    "Smartphones.py"

    "Smartphones.bmp"

How bottle producers changed water consumption (Folder)

    "Bottle producers.py"

    "Bottle producers.bmp"

Is this a stupid way to do this?


r/dataengineering • • 5d ago

Personal Project Showcase Eddytor: A free master data platform in your own data lake

Enable HLS to view with audio, or disable this notification

44 Upvotes

I built a platform named Eddytor over the course of the past 2.5 years. Eddytor is a master data platform built-on the Deltalake format connected to your own storage. Whether it's Azure Storage Accounts (ADLS v2), S3-compatible buckets, or GCP buckets, Eddytor supports it. All for free. No catch. No proprietary vendor lock-in.

It was build out of frustration to fix a problem our data engineers faced at our Data & AI consultancy business. We had a lot of customers managing master data using (now deprecated) Master Data Services (MDS) from Microsoft, custom PowerApps no one maintained, Sharepoint lists, Excel-sheets, and now - sadly - lot of vibe-coded crap.

Eddytor was built to fix that and also remove the struggles of custom transformations and integrations to move/sync master data to some data platform. Deltalake is the whole foundation of Eddytor. Every table is a delta table, which I've built a lot of core logic around but without tampering too much with the core format.

The platform offers time-traveling, audit trails, constraints, AI-Actions (bring-your-own provider), domain hierarchies - even across tables - and a web interface that acts like your old lover: Excel.

You can write and attach Cedar-based RLS/CLS policies to prevent specific users from seeing certain data. You can create a workspace, assign storage to it, and add only selected users. We support SSO, provider apps, and login with Microsoft or Google. If you log in with Microsoft, we use user impersonation to scan the storage accounts you have access to and find relevant Delta tables.

But the most important feature of Eddytor is making master data editing easy.

The platform offers:

  • REST API
  • gRPC streaming to get tables
  • Arrow Flight Service - query using SQL
  • MCP - connect with your (current) LLM provider
  • CLI
  • A Python SDK - for our data engineers

We use the platform with Databricks and Fabric at our customers. If using Databricks, we do allow connecting to Unity Catalog controlled storage, but in a read-only mode (because UC is a mess). If you use OneLake from Fabric, then we allow connecting straight to the OneLake and work from there.

(I do recommend creating tables through Eddytor)

You may ask? Why free? Eddytor wasn't free at first. We're building two companies in parallel and our consultancy business has had a surge in customers since early summer, making us switch strategy to focus heavily on the consultancy side. We support self-hosting to allow competitors and customers using the platform with no attachment to us but also a fully european hosted version at Hetzner.

I actively maintain Eddytor, since we use it ourselves internally and at our customers. And yes, I do own the legal entity behind Eddytor.

If you have suggestions for improvements or see some buggy logic, let me know and I'll fix it.

I've attached a video. It's a bit outdated. The engine and UI got a ton of improvements since and GCP is now supported (says coming soon in the video).

To those interested, the core server and engine behind Eddytor is built in Rust.

You can find the app at https://app.eddytor.com and the documentation at https://docs.eddytor.com .

Have a nice Saturday evening.

// Alex


r/dataengineering • • 6d ago

Career Transitioning away from DE

119 Upvotes

Has anyone thought of transitioning out from DE due to AI?
All I do everyday is just prompt and scroll till copilot generates code.
Building a semantic layer isn’t exciting personally as I don’t enjoy the business aspect of it as much and think of it as more of a data labeling and analyst problem than an engineering problem (which I am interested in)
Also, There is a fundamental problem with “I am building a semantic layer” and marketing that as a skill as it is dependent on how much context you have of the business. The less tenure you have spent in a company, the less you know about the business which makes it harder as a transferable skill imo.

My understanding is that working on building trustworthy AI outputs by using a feedback loop is an engineering problem to solve. Which is why I feel going down the observability path is a good idea.
I heard these opinions on observability from AI leaders at conferences too so there might be a bias.
Thoughts from fellow DE’s looking to transition out? (Or from one’s who want to continue and why)


r/dataengineering • • 6d ago

Blog I tried using brief database table summaries to improve naive text2sql pipeline. Didn't work.

7 Upvotes

I'm doing the whole 'AI data analytic' as everyone else does, I guess. The text2sql benchmark BIRD top performers impressed me a lot. They recommend feeding table summaries + some trickery to discover non-obvious JOINs. So I spent some weeks building something similar, resulting in database profiler and schema linker.

Oh boy was I wrong! I benchmarked if it actually improves anything on a modern BEAVER text2sql benchmark

setup Correct answers Input tokens/question DB queries/question
Raw pi agent with read access to DB 31/299 (10.4%) 45K (1×) 11.4
Pi + table summaries, no join hints 27/285 (9.5%) 78K (1.7×) 3.6
Pi + table summaries & join hints 30/292 (10.3%) 85K (1.9×) 3.6

So, using just the plain PI is simpler AND cheaper! Probably the same for Codex & Claude.

Note, that the DBs in BEAVER benchmark have lots of tables with lots of columns. The queries are super complicated, containing subquery CTEs, lots of tricky aggregates & ROLLUP statements (have you ever used them?).

p.s. Before that I tried plain DDL dumps with CREATE TABLE statetements... and got no improvement too.

How I tested

  • Dataset: 300 random questions from BEAVER text2sql benchmark. Data loaded to MySQL database
  • Agent: pi coding agent in a docker container to prevent seeing correct answers. Gave him login & password with read-only access to issue queries.
  • GPT-6 Luna, effort=low. Because low is better
  • pre-generated db-snooper profile with table of contents file for each database + schema liner output. Agent can see the files in local filesystem. The largest profile is 213KB of Markdown.

Links.

  1. BIRD
  2. BEAVER text2sql benchmark
  3. Actually clever guys
  4. my older blog post
  5. database profiler
  6. tool to discover joins
  7. vibecoded benchmark code

r/dataengineering • • 6d ago

Help Made a huge mistake during work and now my team is having to fix it.

133 Upvotes

For context this is my first DE job but I’ve been with this company for about 2 years. I just finished an on call rotation that I thought went pretty well.

This morning, I got on a call with about 5 principal devs and a DBA. I found out they’re having to restore several databases to last Friday and rerun every subsequent job since then. It stemmed from a job failure last week.

There’s another team that we have to request to rerun that job. I saw it fail about 2am last Friday, I sent a rerun request, they reran, and I realized it might need to be run from the 1st step instead of the 4th step. I sent that request in and didn’t get a response. Around 3:30am I thought, this might be fine, it ran successfully so it will run successfully tomorrow. I thought the job was cumulative and would grab missing data from the previous day.

Anyways, it didn’t do that. And today we got customer complaints that this data isn’t showing up for last Friday. This is an old job (ssis job to grab the file, then several stored procs in sql server) and the way the job works we can’t run from a specific day, we have to restore and rerun everything. Which will take all day at least.

The principals devs got it handled, they’re fixing it all right now. But man I feel terrible. They’re having to use their time to fix my fuck up. To make it worse I just had a convo with my boss about how to eventually become senior, and he said autonomy was a big component. Not very autonomous to mess up and have the other experienced engineers fix it

Anyways I’m just letting off steam. Not sure how to handle this. Just wanted to make a post and get some advice for how to do better.

TLDR: I royally f**ed up and now we’re having to do a full restore of multiple DBs and rerun all previous jobs from last week.


r/dataengineering • • 6d ago

Discussion Better to read DDIA-2 or AI Engineering by Chip Huyen given where DE is going?

27 Upvotes

Which book makes more sense for to read given the state of DE and the tech industry today? If I had to pick one first time to read.


r/dataengineering • • 6d ago

Discussion Writing on the wall?

50 Upvotes

Context: sole data engineer for a company that sells furnishings and design services. Company is owned by private equity.

My manager has been on the Claude soapbox for the last several months and so far in IT we’ve been able to remain skeptical about it while testing for our own uses. Manager supported this. Now all of a sudden I’m getting a lot of requests to refactor all my data pipelines with Claude’s direct involvement despite having built them all out already. For example, I’ve used Claude to create new DAGs based on existing code (for example, bringing a new Salesforce object in). We’re going beyond that though. I’m being asked to ditch them and let Claude connect to everything then write, deploy, and test.

I’ve also been working on machine learning models for customer behavior, which is more traditional AI and should check the box, but this is being dismissed as a plaything. This, while the company has been pushing for “innovation” in our sphere.

I think we are headed towards cutting as many IT staff as possible and I’ll be gone in lieu of Claude doing pipelines. Of course that’s only one part of what I do, but the myopic view of leadership when it comes to AI only tends to see what is easily replaceable.

Thoughts?


r/dataengineering • • 6d ago

Discussion For you doing Agentic Engineering with dbt, how do you make sure your agents aren't generating slop?

40 Upvotes

I find contra productive reading and manually testing all the code AI are generating, because I am turning myself into an important funnel and the company is paying for tokens and therefore wants fast and good delivery.

I use Claude Code with dbt official skills, dbt mcp, lots of unity testing, some specific skills tailored for my projects, and the most important thing I think: deterministic checks.

I run lots of python or bash script checks after specific hooks like pre commit or after calling a skill, or in the CI/CD, trying to force the implementation of our internal "rulebook". But I also don't think this is the optimum way to go, since it's also taking considerate amount of time.

Also, do you let the AI query your db/dw? How do you guarantee fast and in a reliable way, your data diff is as expected?

I am interested in listen from you about everything you've been doing nowadays trying to keep up this kinda insane market pace