MongoDB is a document-oriented NoSQL database that supports high-volume data storage. It organizes data into collections and documents instead of tables and rows. A collection is similar to a table in an RDBMS and holds a set of documents, which are similar to rows. Hevo reads changes from MongoDB through Change Streams.
Hevo supports two ways of configuring MongoDB:
-
MongoDB Atlas: This option is applicable if your MongoDB database is hosted on MongoDB Atlas.
-
Generic MongoDB: This option is applicable for all MongoDB deployments except MongoDB Atlas.
Refer to the required MongoDB variant for the steps to configure it as a Source in your Hevo Pipeline and start ingesting data.
Supported Configurations
| Supportability Category | Supported Values |
|---|---|
| Database versions | 4.0.7 - 8.0 |
| Connection limit per database | No limit |
| Transport Layer Security (TLS) | 1.2 and 1.3 |
Supported Features
| Feature Name | Supported |
|---|---|
| Capture deletes | Yes |
| Custom data (user-configured tables & fields) | Yes |
| Data blocking (skip objects and fields) | Yes |
| Resync (objects and Pipelines) | Yes |
| API configurable | No |
Supported Instance Types
| MongoDB Variant | Instance Types | Supported |
|---|---|---|
| Generic MongoDB | Standalone | No |
| Generic MongoDB | Replica set | Yes |
| Generic MongoDB | Sharded cluster | Yes (only for collections with _id as the shard key) |
| MongoDB Atlas | Replica set | Yes |
| MongoDB Atlas | Sharded cluster | Yes (only for collections with _id as the shard key) |
Prerequisites
-
The MongoDB version is 4.0.7 to 8.0.
-
Use the following command to find out the MongoDB version using
mongosh: .db.version() -
A MongoDB user with the
readAnyDatabaseandclusterMonitorroles on theadmindatabase. Hevo connects to your MongoDB database using password authentication.
Oplog Retention Period
The oplog retention period is the duration for which MongoDB retains entries in the oplog before overwriting them. If Hevo doesn’t read a change before its entry is removed, the Pipeline stops ingesting data, and you must resync the Pipeline. This can happen if:
-
The Pipeline was disabled for longer than the retention period.
-
The oplog filled up, and MongoDB overwrote older entries before Hevo read them.
We recommend sizing the oplog to retain at least 72 hours of changes.
To change the oplog size or retention period, refer to the documentation for your MongoDB variant:
Document Packing Modes
Collections in MongoDB are flexible, which means:
-
Documents can gain new fields.
-
Existing fields can change data types.
-
Documents within the same collection can have different sets of fields.
-
Nested objects can vary in structure across records.
When this data is written to a structured Destination table with fixed columns and data types, these variations can lead to schema changes or replication failures.
Document Packing Modes define how Hevo writes MongoDB documents to the Destination table. These modes determine whether MongoDB fields are flattened into columns or stored as a single JSON document.
You select the packing mode in the Configure Source screen when you configure MongoDB as a Source. The default mode is Unpacked Mode. You cannot change this mode after the Pipeline is created.
Packed Mode
In Packed Mode, Hevo stores each MongoDB document as a single JSON object in a single column of the Destination table.
Hevo does not create separate columns for individual fields. Instead, the document is stored as:
-
Standard metadata columns, such as primary key and ingestion timestamp. The example omits Hevo’s metadata columns.
-
One JSON column that contains the complete document.
For example, consider the following MongoDB document:
{
"_id": ObjectId("652f1c2e9b1e8a3d4c7f0a11"),
"customer": "Anita",
"order_total": 4500,
"address": {
"city": "Pune",
"zip": 411001
}
}
In Packed Mode, the Destination table contains:
| _id (VARCHAR) | document (JSON) |
|---|---|
| 652f1c2e9b1e8a3d4c7f0a11 | {“_id”:”652f1c2e9b1e8a3d4c7f0a11”,”customer”:”Anita”,”order_total”:4500,”address”:{“city”:”Pune”,”zip”:411001}} |
Since the document is stored as JSON, changes to field names, nested objects, arrays, or data types do not modify the Destination table schema.
Packed Mode is suitable when you prefer a stable Destination schema over individual columns. It reduces the chances of failures caused by field additions or type changes and keeps the Destination schema consistent over time.
You can extract individual fields using Models or the JSON functions supported by your Destination.
Unpacked Mode
In Unpacked Mode, Hevo creates one column for each top-level field from each MongoDB document into separate columns in the Destination table.
Using the same example:
{
"_id": ObjectId("652f1c2e9b1e8a3d4c7f0a11"),
"customer": "Anita",
"order_total": 4500,
"address": {
"city": "Pune",
"zip": 411001
}
}
The Destination table contains:
| _id (VARCHAR) | customer (VARCHAR) | order_total (INTEGER) | address (JSON) |
|---|---|---|---|
| 652f1c2e9b1e8a3d4c7f0a11 | Anita | 4500 | {“city”:”Pune”,”zip”:411001} |
Hevo creates columns for top-level fields and infers their data types. If new fields appear or field types change, Hevo updates the Destination schema.
Unpacked Mode is suitable when direct SQL querying, filtering, grouping, and joins on individual fields are required.
Handling of Deletes
Hevo uses MongoDB’s Change Streams to capture data changes, including insert, update, and delete operations. Hevo replicates delete actions captured by the Change Stream to the Destination table by setting the metadata column __hevo__marked_deleted to True for the corresponding row.
Source Considerations
-
The
_idfield in a MongoDB document serves as its primary key. Therefore, Hevo does not ingest a document if its_idfield contains a null value. -
MongoDB does not support BSON documents larger than 16 MB in Change Streams for MongoDB versions earlier than 6.0.9. As a result, Hevo cannot ingest data for such documents.
-
Views and time-series collections are not available for selection as objects.