MongoDB schema design: embed or reference, and how to decide
MongoDB is schemaless in the same way a spreadsheet is: you still have a schema, it is just enforced by whoever wrote the last query.
Choose a relational database when your data has relationships, when you need transactions across entities, or when anyone will ever run an ad-hoc analytical query — which covers most business applications. Choose a document, key-value or wide-column store when the access pattern is narrow and known, documents vary genuinely in shape, or scale demands horizontal partitioning from the start.
| Relational | Document | Key-value | Wide-column | |
|---|---|---|---|---|
| Example | PostgreSQL, MySQL | MongoDB | Redis, DynamoDB | Cassandra |
| Best access | Any query, including ad hoc | By key or indexed field | By key | By partition key |
| Joins | Native and efficient | Limited, expensive | None | None |
| Transactions | Full ACID | Multi-doc with cost | Limited | Limited |
| Schema change | Migration | Application-managed | N/A | Flexible columns |
| Horizontal scale | Harder, improving | Built for sharding | Trivial | Built for scale |
When entities relate to each other and you cannot predict every query in advance.
The 'relational databases do not scale' argument is largely historical for typical product workloads. A well-indexed PostgreSQL instance on modern hardware handles tens of thousands of transactions per second and hundreds of gigabytes without exotic architecture.
A typical production system we build uses PostgreSQL as the source of truth, Redis for caching and rate limiting, object storage for files, and a dedicated search index when full-text needs outgrow the database. Each has a clear job and a clear owner.
The cost is operational: more backups, more monitoring, more failure modes, more expertise. Add a second database when a specific workload demands it, not because a diagram looks more sophisticated with it.
Yes, and most mature systems do. Keep one clear source of truth and use other stores for specific workloads — caching, search, analytics — rather than splitting ownership of the same data.
For its designed access pattern — a single-key lookup on a partitioned dataset — yes. For anything requiring joins, aggregation or ad-hoc filtering, a relational database is usually faster because it was built for exactly that.
Distributed SQL systems offer horizontal scale with relational semantics and are a strong option at large scale. They add operational and latency considerations, so evaluate against your real workload rather than the marketing.
Hard, and it gets harder with every month of data and every query written against the old model. This is exactly why the decision is worth an hour of structured thought at the start.
Harshal Patel
Founder & Lead Engineer, ROVQIX
Harshal leads engineering at ROVQIX, where he has shipped production Next.js, Node.js and PostgreSQL systems for startups, SaaS teams and ecommerce brands. He writes about the trade-offs behind architecture decisions rather than the framework of the week.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
MongoDB is schemaless in the same way a spreadsheet is: you still have a schema, it is just enforced by whoever wrote the last query.
Both are excellent. The differences that matter now are extensibility, data type strictness and operational culture, not raw speed.
Read Committed prevents dirty reads. It does not prevent two people spending the same balance — and that is the bug you will get.
No spam. Just the occasional case study and craft breakdown.