Most AI agents built on Node.js and MongoDB store their memory the same way: one document per session, with the conversation in a messages array. It works in a demo. In production it loses turns when two requests overlap, writes the same answer twice when a job retries, and keeps growing until it hits a hard limit.
This article shows each failure with code you can run, then replaces the design with one message per document, deterministic keys and a per-session lease. The examples use the official mongodb driver and BullMQ.
The design most agents start with
The examples are ESM TypeScript run with tsx, against a local MongoDB and Redis.
npm install mongodb bullmqA stand-in for the model keeps everything runnable without an API key. The half-second delay matters: it is the window in which things go wrong.
export type Role = "user" | "assistant" | "tool";export interface ChatMessage { role: Role; content: string;} // Replace with your model client. The delay stands in for a real model call.export async function callModel(context: ChatMessage[]): Promise<string> { await new Promise((resolve) => setTimeout(resolve, 500)); return `You said: ${context.at(-1)?.content ?? ""}`;}import { MongoClient, ObjectId } from "mongodb";import type { Role } from "./model.js"; export interface Session { _id: ObjectId; userId: string; nextSeq: number; lockedUntil: Date | null; lockedBy: string | null; createdAt: Date;} export interface Message { sessionId: ObjectId; seq: number; key: string; role: Role; content: string; createdAt: Date;} const client = new MongoClient(process.env.MONGODB_URI ?? "mongodb://localhost:27017");await client.connect(); export const db = client.db("agent");export const sessions = db.collection<Session>("sessions");export const messages = db.collection<Message>("messages"); await messages.createIndex({ sessionId: 1, key: 1 }, { unique: true });await messages.createIndex({ sessionId: 1, seq: -1 });The Session and Message types are for the fixed design later on. The naive version keeps everything in one document:
import type { ObjectId } from "mongodb";import { db } from "./db.js";import { callModel, type ChatMessage } from "./model.js"; interface NaiveSession { _id: ObjectId; messages: ChatMessage[];} const naive = db.collection<NaiveSession>("naive_sessions"); export async function runTurn(sessionId: ObjectId, text: string) { const session = await naive.findOne({ _id: sessionId }); if (!session) throw new Error("no such session"); const history: ChatMessage[] = [...session.messages, { role: "user", content: text }]; const reply = await callModel(history); await naive.updateOne( { _id: sessionId }, { $set: { messages: [...history, { role: "assistant", content: reply }] } }, );}Read the history, call the model, write the history back. Agent loops are shaped like this because the framework holds the conversation in memory for the length of a turn, and saving it at the end is the obvious place to persist.
Two turns at once: one of them disappears
Send two messages to the same session before the first reply lands. A double-clicked send button, a second tab or a client retry on a slow response is enough.
import { ObjectId } from "mongodb";import { db } from "./db.js";import { runTurn } from "./naive.js"; const _id = new ObjectId();await db.collection("naive_sessions").insertOne({ _id, messages: [] }); await Promise.all([runTurn(_id, "first"), runTurn(_id, "second")]); const after = await db.collection("naive_sessions").findOne({ _id });console.log(after?.messages.length);process.exit(0);$ npx tsx src/race.ts2Four messages were produced; two were kept. Both turns read an empty history, both called the model, and each wrote back its own version of the array. The second $set replaced the first. Nothing errors, nothing is logged, and the user sees a reply that is missing from the history on the next turn.
Switching to $push fixes this particular lost update, because the server appends atomically. It does not fix the next two problems.
A retried job answers twice
Put the turn on a queue so it survives a restart, and the worker will one day crash between writing the reply and acknowledging the job. The queue does what it should and runs the job again. With $push, the retry calls the model a second time and appends a second, different answer to the same question.
This is not rare. Anything between "the write happened" and "the job was acked" — a deploy, an out-of-memory kill, a Redis blip, a stalled-job check that fires during a long model call — produces it.
A retry that runs the model again is not a retry; it is a second answer.
The fix is to name every message after the work that produced it, so the database can recognise a repeat.
The session document only grows
Every turn adds to the array and nothing takes away. That has two costs.
The hard one: a BSON document cannot exceed 16 mebibytes, according to MongoDB's limits reference. Agents that store tool output — search results, fetched pages, file contents — in the history reach it far sooner than a chat of short messages would. When they do, the write fails and that session is stuck.
The soft one comes first: every turn reads the whole history to use the last few dozen messages. The document gets slower to load long before it gets too big to write.
One document per message
The replacement keeps a small session document for the counters and one document per message.
| Collection | Holds | Written by |
|---|---|---|
sessions | nextSeq, the lease fields, the owner | $inc and conditional updateOne |
messages | one message, its seq and a deterministic key | insertOne only |
Messages are never updated. A new message is an insert; a repeated one is rejected by the unique index.
Keys that make a repeat detectable
Each message gets a key that depends only on the work, never on when it ran: <turnId>:user, <turnId>:assistant, <turnId>:tool:<callId>. The turnId comes from the client, generated once per message the user sends, so a double-submitted request carries the same id.
The unique index on { sessionId: 1, key: 1 } in src/db.ts turns "did we already write this?" into a question the database answers atomically. The second index, { sessionId: 1, seq: -1 }, serves the context query below.
Appending without losing or doubling
import { MongoServerError, ObjectId } from "mongodb";import { messages, sessions } from "./db.js";import type { ChatMessage, Role } from "./model.js"; export async function createSession(userId: string): Promise<ObjectId> { const _id = new ObjectId(); await sessions.insertOne({ _id, userId, nextSeq: 0, lockedUntil: null, lockedBy: null, createdAt: new Date(), }); return _id;} export async function appendMessage( sessionId: ObjectId, key: string, role: Role, content: string,): Promise<"inserted" | "duplicate"> { const session = await sessions.findOneAndUpdate( { _id: sessionId }, { $inc: { nextSeq: 1 } }, { returnDocument: "before", projection: { nextSeq: 1 } }, ); if (!session) throw new Error(`session ${sessionId.toHexString()} not found`); try { await messages.insertOne({ sessionId, seq: session.nextSeq, key, role, content, createdAt: new Date(), }); return "inserted"; } catch (err) { if (err instanceof MongoServerError && err.code === 11000) return "duplicate"; throw err; }} export async function hasMessage(sessionId: ObjectId, key: string): Promise<boolean> { return (await messages.countDocuments({ sessionId, key }, { limit: 1 })) > 0;}$inc hands out sequence numbers atomically, so two concurrent appends never share one. A duplicate key error (code 11000) means the message is already there, which is success, not failure. A duplicate burns a sequence number and leaves a gap; gaps are harmless because nothing counts them.
Reading the window on turn boundaries
export async function loadContext(sessionId: ObjectId, limit = 40): Promise<ChatMessage[]> { const recent = await messages .find({ sessionId }, { projection: { _id: 0, role: 1, content: 1 } }) .sort({ seq: -1 }) .limit(limit) .toArray(); recent.reverse(); const start = recent.findIndex((m) => m.role === "user"); return start === -1 ? [] : recent.slice(start);}The query walks the { sessionId: 1, seq: -1 } index backwards and stops at limit, however long the session is. The last two lines matter as much as the query: a window cut at a fixed count can start in the middle of a turn, handing the model a tool result without the call that produced it. Dropping everything before the first user message keeps whole turns only.
One turn per session at a time
The unique keys stop duplicates. They do not stop two different turns from running side by side, each building its context without the other's messages. For that, a turn takes a lease on its session before it starts.
import type { ObjectId } from "mongodb";import { sessions } from "./db.js"; export async function acquireLease(sessionId: ObjectId, owner: string, ms: number) { const now = new Date(); const result = await sessions.updateOne( { _id: sessionId, $or: [{ lockedUntil: null }, { lockedUntil: { $lt: now } }] }, { $set: { lockedUntil: new Date(now.getTime() + ms), lockedBy: owner } }, ); return result.modifiedCount === 1;} export async function releaseLease(sessionId: ObjectId, owner: string) { await sessions.updateOne( { _id: sessionId, lockedBy: owner }, { $set: { lockedUntil: null, lockedBy: null } }, );}The filter and the update are one operation, so only one caller can move lockedUntil from "free or expired" to "taken". Release matches on lockedBy, so a worker whose lease already expired cannot release someone else's.
A lease serialises turns; it does not order them. If turn B takes the lease before turn A, B runs first. Chat clients rarely send B before A has an answer, but if yours can, order has to come from the client.
The worker that ties it together
import { Queue } from "bullmq"; export interface TurnJob { sessionId: string; turnId: string; text: string;} const turns = new Queue<TurnJob>("turns", { connection: { host: "localhost", port: 6379 } }); export async function enqueueTurn(job: TurnJob) { await turns.add("turn", job, { jobId: `turn-${job.sessionId}-${job.turnId}`, attempts: 8, backoff: { type: "exponential", delay: 2000 }, removeOnComplete: { age: 24 * 3600 }, });}The job id is built from the same turnId, so a replayed request is dropped at the queue while the first job is still stored. BullMQ's job id guide is explicit about two limits: once a job is removed — here, a day after it completes — the same id can be added again, and custom ids must not contain :. The first limit is why the MongoDB keys are still needed; the second is why the id uses dashes.
import { Worker } from "bullmq";import { ObjectId } from "mongodb";import type { TurnJob } from "./enqueue.js";import { acquireLease, releaseLease } from "./lease.js";import { appendMessage, hasMessage, loadContext } from "./memory.js";import { callModel } from "./model.js"; const LEASE_MS = 120_000; new Worker<TurnJob>( "turns", async (job) => { const sessionId = new ObjectId(job.data.sessionId); const { turnId, text } = job.data; const owner = `${job.id}.${job.attemptsMade}`; if (!(await acquireLease(sessionId, owner, LEASE_MS))) { throw new Error(`session ${job.data.sessionId} is busy`); } try { await appendMessage(sessionId, `${turnId}:user`, "user", text); if (await hasMessage(sessionId, `${turnId}:assistant`)) return; const reply = await callModel(await loadContext(sessionId)); await appendMessage(sessionId, `${turnId}:assistant`, "assistant", reply); } finally { await releaseLease(sessionId, owner); } }, { connection: { host: "localhost", port: 6379 }, concurrency: 10 },);Walk through the crash from earlier. The worker wrote the assistant message and died before the ack. On the retry, the user message comes back as a duplicate, hasMessage finds the answer, and the job returns without calling the model. No second answer, no second bill.
A busy session throws, so BullMQ retries it with backoff. Each of those retries uses an attempt, which is why attempts is higher than you would set for a job that only fails on real errors.
Giving an agent query access without giving it the database
Memory is what the agent writes. The other half is what it reads: sooner or later someone gives the agent a tool that queries MongoDB with a filter the model wrote.
For coding agents working against Atlas, MongoDB has moved this into the platform. Atlas App Connections, announced with a hosted Atlas MCP server, replaces shared service accounts with OAuth 2.1 delegation: the tool acts with the developer's own permissions, AI client access is off until an admin enables it, and admins can force read-only mode, set token lifetimes and revoke access.
Agents inside your own product get none of that. The boundary is yours to build, and three rules cover most of it: run the tool as a database user with only the read role on the data it needs, take the tenant from the authenticated session rather than from the model, and allowlist the query operators the model may use.
import { db } from "../db.js"; const ALLOWED = new Set(["$eq", "$ne", "$in", "$gt", "$gte", "$lt", "$lte", "$and", "$or"]); export function assertSafeFilter(value: unknown): void { if (Array.isArray(value)) return value.forEach(assertSafeFilter); if (value === null || typeof value !== "object") return; for (const [key, inner] of Object.entries(value)) { if (key.startsWith("$") && !ALLOWED.has(key)) { throw new Error(`operator ${key} is not allowed`); } assertSafeFilter(inner); }} export async function findOrders(userId: string, filter: Record<string, unknown>) { assertSafeFilter(filter); return db .collection("orders") .find({ $and: [filter, { userId }] }, { projection: { _id: 0, number: 1, status: 1, total: 1 } }) .limit(50) .toArray();}An allowlist rather than a blocklist, because the operators you would forget to block — $where, $function, $expr — are the ones that let a filter do more than filter. The $and with userId means the model can narrow the results but never widen them to another customer's orders.
Before you ship
- No code path reads the history, changes it in memory and writes the whole array back.
- Every message has a key derived from the turn, and a unique index enforces it.
- A duplicate key error on append is treated as success.
- The worker checks for an existing answer before calling the model.
- One turn per session holds a lease longer than the slowest model call.
- Context is read from an index, with a limit, cut on turn boundaries.
- Tools that let the model query the database run read-only, scoped to the caller, with an operator allowlist.
If you fix only one thing, make it the keys. The lease and the queue settings reduce how often things collide; the unique index is what makes a collision harmless.