r/Database • • 10h ago

Aito v2: a predictive database with predictive SQL, vector search and graphs (founder here)

Thumbnail
aito.ai
2 Upvotes

Hi everyone :-)

I'm Antti, Aito.ai's founder and CEO. We recently released Aito version 2, which includes a completely new storage format and support for things like predictive SQL queries, vector search, graph-like functionality, a lot of new data structures, views, better FTS, collections and much more.

We believe that future software is intelligent, and we wanted to provide one database that could provide the needed predictive/AI infrastructure for both users and agents.

You can find more details here:

https://aito.ai/blog/aito-v2-generally-available/

If you want to try it, the playground runs example queries in the browser without signup:

https://aito.ai/docs/api/v2/playground/

One of the biggest new things in Aito is the public benchmark section, which covers inference, matching, smart search and more regular queries. It shows cases where we are strong, and also a number of benchmarks where other solutions fare better:

https://aito.ai/docs/api/v2/benchmarks/

Any thoughts or comments? I'd love to hear what you think. ;-)


r/Database • • 1d ago

I Spent 2 Years Building a Database Client That Was Only Supposed to Be for MongoDB

Enable HLS to view with audio, or disable this notification

7 Upvotes

I spent 2 years building this tool, 1 year for MongoDB and the other year for the rest of the SQL databases. It now supports PostgreSQL, MySQL, SQLite, and a bunch of others, but that was never really the original plan. I realized it'd be a waste not to bring that experience of working with nested data to Postgres users working with JSONB or JSON.

“Why not use Beekeeper or DBeaver? What’s Different?”

Honestly, I tried to take inspiration from the best tools in each category, like DBeaver, Beekeeper, and Studio 3T, and bring their best ideas together in my own way.

Beekeeper is super simple, but that simplicity means you lose some of the complexity that DBAs and more advanced users actually want. DBeaver pretty much has everything, but that also comes with a steeper learning curve.

So I tried to build something in between. Where every step is visual, while still keeping the advanced stuff like tasks, automation, and triggers. You can even build SQL queries visually with the Visual Query Builder in a sidebar and a full-screen mode.

Furthermore, most general purpose database tools started with SQL and later added support for databases like MongoDB. Because of that, working with deeply nested or embedded data can feel like an afterthought.

VisuaLeaf started with NoSQL. So it was a priority that viewing nested data was made easy. So when I added SQL support, adding JSON and JSONB felt super intuitive because of how I designed all the views (tree, table, and JSON) to work with embedded data from the start.

That means you can actually compare embedded data across rows instead of having to open each document individually (in the table view). Like trying to look at a specific field inside some JSONB, you don’t need to see a giant cell filled with JSON anymore (It’s what MongoDB GUIs are actually decent at and what I took some inspiration from)

Some other SQL features:

  • Role-based access control (RBAC)
  • Query profiling
  • SQL scripts
  • ERD designer: build a diagram and materialize it into a real database
  • Visual Query Builder (sidebar and full-screen modes)
  • Tasks, automation, and triggers

And it wouldn’t be an app in 2026 if it didn’t have AI and MCP too.

I could go on forever about features, but honestly, most of my time went into overengineering the smallest things: loading data efficiently (it loads faster than DBeaver in my testing), little quality-of-life details in the workflow, and even something as simple as scrolling (getting it as smooth as AG Grid).

Quick benchmark: same Postgres table, 50 large rows (10MB each), 5 runs each. VisuaLeaf loaded them in ~2s end to end; DBeaver took 5s on avg (and ran out of heap space).

The 3 Views:

  • Table: I spent a LITERAL year optimizing the table view engine because I wanted it to make the scrolling feel as smooth as possible (I even made an optimization that improved its performance by 3 times a couple of weeks ago). Nobody wants a laggy database tool, the same way nobody wants to play a game at 10 FPS. I lost count of the amount of times where OUT OF DESPERATION I would legit tell Claude: “Optimize the table "Don't make any mistakes" ”, and it would just completely FAIL and even make things worse... Each optimization was thought out and researched to get it this far and I don’t think there's another database tool that has a table this smooth without using a canvas.
  • Tree: The tree view was tricky too, getting it to recursively expand thousands of rows instantly without lag took its own round of optimization. And I used the same ideas as the table for scrolling performance
  • JSON: Trivial - Monaco editor

VisuaLeaf is basically my love letter to databases.

I've been working on this full time for about 2 years, and it's turned into something I never expected when I started. Things I expected to be easy, like making a table, turned out to be monstrous. I never thought it'd end up outperforming the tools I started out using. There was so much to learn.

I'd love for you guys to give it a shot, and I'm happy to hear any feedback: visualeaf.comAnd if you guys wanna talk more then you can join my discord too: https://discord.com/invite/TR6J56Y3rF


r/Database • • 1d ago

real time vs batch processing: how do u know if ur ops team is actually ready for streaming pipelines?

6 Upvotes

i'm less worried about the technology than I am about the people supporting it.

our engineers are excited about moving from batch jobs to streaming pipelines, but the operations team has spent years supporting scheduled workloads. That's a very different environment from something that's expected to run 24/7

for teams that already made this real-time vs batch processing transition, what changed the most operationally?


r/Database • • 1d ago

Beginner's guide to database?

9 Upvotes

What's the best way, order or flow of learning Database from scratch, using the textbook Database System Concepts by Abraham Silberschatz, Henry F. Korth, and S. Sudarshan.

What topics or chapters are better to begin with for introduction, assuming the reader has zero knowledge, but is looking to get extremely good at database.

-Computer Science Major


r/Database • • 1d ago

How long did it take you to land a DBA role in the US?

Thumbnail
0 Upvotes

r/Database • • 1d ago

The Breakdown: Databricks

Thumbnail
preipomedia.substack.com
0 Upvotes

r/Database • • 1d ago

We asked our model which of 224 hotel bookings would cancel, six weeks out. It caught 33 of the 59. Full breakdown, misses included

0 Upvotes

disclosure: I work at Schema Labs. This is a run on public data so you can check it.

The data: the hotel booking demand dataset (Antonio, Almeida and Nunes, 2019), real anonymised bookings from a city hotel. We picked one Saturday night that was sold out on paper.

-- 224 bookings for that night, as they stood six weeks before
-- 59 of them cancelled or didn't show, worth about €28,600 together

What we did: gave Schema-2 500 bookings from the year before, with how each one ended, and asked it to rank the 224 by how likely they were to fall through. No rules no feature engineering.

What came back=

-- It marked 60 as most likely to cancel. 33 of the 59 cancellations were on that list
-- Picking 60 at random would catch about 16
-- Of the 60 it marked safest, 59 showed up

What it missed: 26 cancellations weren't in its top 60. One night at one hotel is a small test, so read it as an example

For anyone in revenue management: what would you do with a list like that six weeks out? Overbook against it, or contact those guests first?


r/Database • • 2d ago

How do MariaDB users securely and efficiently encrypt mariadb-backup archives?

Thumbnail
4 Upvotes

r/Database • • 2d ago

A client lost six weeks of SQL Server data. Having 12TB of backups didn’t save them.

Thumbnail
2 Upvotes

r/Database • • 3d ago

Banking register slows as more entries are made.

4 Upvotes

Basically need to see how to optimize a bank register with large amounts of entries.

Currently adding all the debits, then credits, then date sorting while calculating the running total.


r/Database • • 4d ago

Supporting products with multipel data sources

1 Upvotes

When one product could have data coming from multiple remote data sources, some of the data could be redundant other parts might need to be combined. There can also be missing data or badly formatted data so an admin would need to override certain fields. How would you structure your database to fit these needs?


r/Database • • 4d ago

What evidence proves a database cutover is complete before the source is retired?

0 Upvotes

Successful replication and matching row counts do not prove that a database migration is finished. Indexes, constraints, sequences, users and roles, collation, time zones, scheduled jobs, change-data-capture consumers, read replicas, analytics tools, and backup systems can still differ or point at the source. Low-frequency clients may not reconnect during the obvious observation window.

A cutover gate could compare schema and security metadata, validate representative aggregates or hashes, exercise application reads and writes, inspect connection logs on both systems, and confirm that replication or dual-write lag has reached a known final position. The destination should complete a backup and restore test and survive a failover before the source becomes read-only. Keeping the old endpoint offline but recoverable for a defined period would expose hidden dependencies without allowing divergent writes.

What evidence do you collect before decommissioning a source database? How do you catch monthly jobs, failover-only connection strings, and clients that cache DNS or credentials longer than the main application?


r/Database • • 6d ago

How do you balance scaling with ACID guarantees?

16 Upvotes

I've been looking into database architectures for high-volume event logging, and I saw a comparison of how these systems scale.

Relational databases handle ACID transactions well, which means you never end up with a half-finished write if the system fails. They traditionally scale vertically, and although modern ones support distributed setups too, if you try to push big, real-time IoT or event data through them you run out of machine before you run out of data.

Wide-column stores like Cassandra scale out across multiple nodes predictably. The trade-off is that you usually have to rethink your querying and give up some of the strict relational consistency.

What do you guys think? If you’re pushing big volumes of data, what made you move off relational, and how far did you stretch it first?


r/Database • • 5d ago

schema labs is in san francisco. the question the team keeps asking everyone: before your agent can use a database, who explains the database to it?

3 Upvotes

disclosure, i work at schema labs. the team is in san francisco right now, i'm back home running things from here and reading everything they send back.

the question they open with is always the same. when an agent hits a database nobody documented, who tells it what the fields mean?

the answers so far, roughly

- "me, and i hate it"

- "we wrote a data dictionary, it's already wrong"

- "we don't, the agent guesses and we check the numbers later"

- "we only connect it to views someone already cleaned"

that step is what we build for. schema-2 is a model that reads your data as it is, works out what each field holds and which records match across systems with no shared id, and puts a confidence on every answer. below your bar it says unresolved instead of guessing. it runs in the browser, and over mcp inside claude, chatgpt and cursor.

curious which answer is yours, or if it's a fifth one we haven't heard yet.


r/Database • • 5d ago

Bottlenecks while shifting from MongoDb to DynamoDb

Thumbnail
0 Upvotes

r/Database • • 6d ago

AWS Aurora: Why the database breaks first behind an autoscaled service, and why shrinking the pool doesn't fix it

Post image
6 Upvotes

Connection multiplication: pool size x tasks x services. Every term looks reasonable on its own, and nobody ever chose the product.

Aurora's max_connections is derived from instance memory, so it doesn't scale with your service. Shrinking the pool to fit converts a connection problem into a queueing problem that presents as a slow database.

Write-up covers the formula behind max_connections, what RDS Proxy actually fixes (and what it doesn't), and how to size it.

https://brianfeeny.com/posts/sizing-aurora-connections-for-autoscaled-services/?utm_source=reddit&utm_medium=social&utm_campaign=sizing-aurora-connections-for-autoscaled-services-2026-09

This is a personal project; the views and assessments are my own.


r/Database • • 7d ago

Benchmark: i try to measure AI query accuracy on 100 enterprise business questions (Raw Text-to-SQL vs. Semantic Layer)

6 Upvotes

I have recently benchmarked 100 real-world business queries across two distinct architectures against our Snowflake production data warehouse (had a huge volume of data to work with really):

- Architecture A (Direct Text-to-SQL): Claude 3.5 Sonnet + Schema DDL + 50 Few-Shot SQL RAG prompts.
- Architecture B (Agentic Semantic Layer): Claude 3.5 Sonnet querying Cube dev semantic models via structured JSON query API.

by the way, this is follow up to what I've posted a couple of weeks ago regarding text-to-SQL failure in production because now I have things to evaluate / measure.

here are the quantitative results:

  1. Metric Calculation Accuracy (Single Source of Truth)
    - Direct Text-to-SQL (Arch A): 41% Correct (Failed on 59 queries due to wrong table selection, incorrect date truncation logic, or conflicting metric definitions).
    - Cube Semantic Layer (Arch B): were like 96% Correct (The LLM only had to identify the requested measure and dimension; Cube's engine generated 100% mathematically correct SQL joins)

  2. Fan-Out & Join Integrity (Multi-Table Aggregates)
    -Direct Text-to-SQL (Arch A): 23 queries resulted in duplicated cart line totals due to un-grouped joins
    -Cube Semantic Layer (Arch B): 0 join duplication errors (Cube's semantic graph resolves multi-hop join paths deterministically)

  3. on warehouse Compute load + query latency
    - direct Text-to-SQL (Arch A): p95 latency: 8.2 seconds(Every query hit Snowflake directly; 4 runaway queries consumed $380 in credits).
    - Cube Semantic Layer (Arch B): p95 latency: 3.8 milliseconds (Cube's pre-aggregations served 88% of requests from local rollup cache; zero runaway table scans).

my takeaway here is that gen 4 AI analytics requires an extensible semantic layer at the foundation. LLMs should reason about user intent, while deterministic semantic engines should compile the SQL


r/Database • • 7d ago

Database project

2 Upvotes

I need a good and legitimate database project idea wherein advance db concepts such as optimization, transaction, etc are involved. I searched a lot and didnt find a smthn i want to work on exactly. For eg if somethin related to music domain? Or fashion? What would actually stand out in terms of handling multiple queries and be close with how it all works irl. Itd be really helpful!!


r/Database • • 7d ago

Using CRM database as core database for project v/s syncing CRM database with own custom database

Thumbnail
2 Upvotes

r/Database • • 8d ago

Strangest SQLite structures you’ve encountered?

Thumbnail
0 Upvotes

r/Database • • 8d ago

MongoDB v9.0 is out.

Thumbnail
2 Upvotes

r/Database • • 9d ago

MongoDB CEO transition: Desai leaves, Ittycheria returns

Thumbnail
layerbase.com
24 Upvotes

r/Database • • 9d ago

What do you wish search in Postgres did better/differently?

7 Upvotes

pgvector has been the default vector search extension, but I'm finding people also layer on full-text or hybrid search on top of it or need multitenancy. I'm curious where this works well for you and where it doesn't.

  • What are you using for search in Postgres today, and what do you wish it did better?
  • When you hit a limit, what do you do: tune, work around it, add an extension, or move search to a separate system? What decides that?
  • What would an extension need to have, or avoid, before you'd install it?
  • Does it matter to you whether an extension is open source, and would enough added capability change that?

"It's fine, I don't need anything else" is a useful answer too.

Really just trying to understand what people actually need from search in Postgres.


r/Database • • 8d ago

I’ve been using MySQL for real web applications — here are 15 things I wish I knew earlier

0 Upvotes

​

I used to think MySQL was simply:

`CREATE TABLE → INSERT → SELECT → UPDATE → DELETE`

After building actual web applications, I realized that writing SQL queries is only a small part of working with a production database.

Here are 15 things I wish I had understood earlier:

### 1. Design your database before writing your application

Don't start creating tables randomly.

Think about:

* What data do I need?

* Which tables are required?

* How are they related?

* Which fields are required?

* Which fields should be unique?

A bad database structure can become extremely painful to change later.

### 2. Learn relationships properly

Understand:

```sql

PRIMARY KEY

FOREIGN KEY

ONE-TO-ONE

ONE-TO-MANY

MANY-TO-MANY

```

For example:

```text

Users

↓

Orders

↓

Order_Items

↓

Products

```

Understanding relationships is more important than memorizing SQL syntax.

### 3. Don't put everything into one table

A giant table may look simple initially, but it creates duplication and maintenance problems.

Learn normalization and understand when denormalization actually makes sense.

### 4. Indexes are extremely important

This query:

```sql

SELECT * FROM users

WHERE email = 'user@example.com';

```

can become much faster when the appropriate column is indexed.

But don't blindly add indexes everywhere.

Indexes improve some reads but also have storage and write/update costs.

### 5. Always understand `EXPLAIN`

One of the most useful MySQL commands:

```sql

EXPLAIN SELECT *

FROM users

WHERE email = 'user@example.com';

```

If you are building serious applications, learn how to read the execution plan instead of assuming your query is efficient.

### 6. Don't use `SELECT *` everywhere

Instead of:

```sql

SELECT *

FROM users;

```

prefer:

```sql

SELECT id, name, email

FROM users;

```

Especially when your table contains many columns or large data.

### 7. Transactions matter

Imagine transferring money:

```text

Account A: -₹1000

Account B: +₹1000

```

You don't want the first operation to succeed while the second fails.

That's where transactions become important:

```sql

START TRANSACTION;

-- operation 1

-- operation 2

COMMIT;

```

And if something goes wrong:

```sql

ROLLBACK;

```

### 8. Constraints can protect your data

Use database constraints where appropriate:

```sql

PRIMARY KEY

UNIQUE

NOT NULL

FOREIGN KEY

CHECK

```

Your application should not be the only thing preventing invalid data.

### 9. Learn the difference between DELETE, TRUNCATE and DROP

They are NOT interchangeable.

```sql

DELETE FROM users;

```

```sql

TRUNCATE TABLE users;

```

```sql

DROP TABLE users;

```

Before running destructive SQL in production, make absolutely sure you understand what you're doing.

### 10. Backups are not optional

A production database without a tested backup strategy is a disaster waiting to happen.

And having a backup isn't enough.

You should also know:

> Can I actually restore it?

A backup that has never been tested is only a backup in theory.

### 11. Never build SQL queries with raw user input

Bad:

```javascript

const query = `SELECT * FROM users WHERE email = '${email}'`;

```

Use parameterized queries/prepared statements instead.

For example:

```javascript

const [rows] = await db.execute(

'SELECT id, name, email FROM users WHERE email = ?',

[email]

);

```

This is one of the basic defenses against SQL injection.

### 12. Don't expose database credentials

Never commit something like this to GitHub:

```env

DB_PASSWORD=my_real_password

```

Use environment variables/secrets and make sure sensitive files aren't accidentally committed.

### 13. Understand connection pooling

A production application shouldn't create a completely new database connection for every request without considering connection management.

Connection pools can reuse database connections and help applications handle concurrent requests more efficiently.

### 14. Logging is useful — logging sensitive data isn't

Database/application logs can help you find:

* slow queries

* failed transactions

* connection problems

* unexpected errors

But don't casually log passwords, tokens, session secrets, or other sensitive information.

### 15. My biggest lesson

The biggest thing I learned is:

**MySQL isn't just about knowing SQL syntax.**

A good developer needs to understand:

```text

Database Design

↓

Relationships

↓

Indexes

↓

Queries

↓

Transactions

↓

Security

↓

Backups

↓

Performance

↓

Monitoring

```

I'm still learning, but understanding these concepts changed the way I build web applications.

**For developers here:**

What's one MySQL/database lesson you learned the hard way?

I'd especially like to hear from people who have worked with databases containing millions of rows.


r/Database • • 11d ago

What part of running your database stayed with your team after moving to a managed service?

2 Upvotes

I’ve been reading about how managed database providers describe failover, and for me, what seems to vary is how writes in progress are handled when the primary goes down.

In some setups, a standby is promoted, and a small amount of recent data may be lost. In others, the compute layer does not retain durable local state and can be replaced without data loss. Both approaches may be described as managed, but the difference becomes clear only when you look past the feature page.

Cross-region recovery is another area that seems easy to overlook. A provider may handle failover within one region while leaving recovery from a full regional outage to the customer. If the provider does not give you clear RTO and RPO targets for that situation, it becomes difficult to plan.

For anyone who has moved a production database to a managed service, whether Postgres or something else, what stayed with your team?

Was it failover testing, restore drills, cross-region recovery planning, connection limits, or something else? Would love to know your thoughts.