Since our Copilot product exited beta in Q2 2025, we've been building AI agents to help accounting professionals automate their books. Last month we released our most ambitious agent yet, the Accounting Specialist, which helps accountants efficiently apply journal entries to similar groups of unclassified transactions. The agent is easy to set up in a Copilot chat session or in the Agent Marketplace, and once activated it begins autonomously grouping unclassified transactions, creating unique cohorts across many dimensions to represent recurring activities. The agents notify customers when new transaction groups require human input for classification, deeplinking into a guided review flow for applying journal entry templates in bulk and resolving one-off exceptions that fall outside of logical grouping patterns.
The Accounting Specialist was one of the most requested agents by our customers and we iterated through multiple rounds of UX to design an intuitive, efficient user experience (and of course, we are continuously seeking feedback to improve it!). In development, we were challenged with evolving the capabilities of our multi-agent platform and establishing new UI conventions to achieve the target design. In this first of two posts, we share some of the challenges and lessons learned with context engineering while developing this agentic workflow.

Accounting Specialist presents a summary of unclassified transactions and review groups
Context Window Management
Context window management has emerged as an industry-wide challenge for any company building AI-driven chat products. Essentially, every large language model (LLM) has a limited buffer ("window") of data that is sent with a chat generation request. As a conversation naturally grows in size with new messages (including generated tool outputs), there is a critical need to continuously prune (effectively "curate") the new context sent with each subsequent request (i.e. conversation turn). While a chat app should always be prepared to handle errors from LLM APIs, an overflow resulting from an unmanaged context window can completely block a conversation from continuing.
A core engineering (and product!) challenge we've faced while building our Copilot chat is how to manage this context, namely for fulfilling customer requests involving large data sets. This limitation has been apparent since we built our first AI chat product and has required an evolving set of techniques and practices to ensure our agents have relevant information to serve user-stated objectives.
Token budgets
LLMs across model providers differ in design. Each encodes and processes information differently according to its training data, architecture, and optimization trade-offs. Context window capacity is measured in tokens rather than bytes, with tokenization schemes varying by provider. A token is the fundamental unit of input, corresponding to a word, subword, character, or punctuation mark depending on the tokenizer. In multimodal models, non-text inputs such as images or audio are also represented as tokens, for example as image patches, spectrogram segments, or frame embeddings.
State-of-the-art frontier models in 2025 range from ~32k tokens (smaller, faster models) up to 1M tokens (larger flagship models). While one million tokens can represent roughly 700k–800k words (comparable to about 1.5-1.7× the length of The Count of Monte Cristo) it is still small relative to enterprise data volumes. Supersized contexts, such as Magic's 100M context window, extend this further but still cap out at only hundreds of megabytes (rough approximation), far below the scale of typical corporate datasets. The key point is that no matter how large the window, it only covers bounded slices of data, making effective context management essential.
AI Development Framework Selection
AI development tooling has grown rapidly in the last two years, with many new frameworks emerging to help software teams manage prompts, memory, retrieval, and model orchestration. When we set out to build Copilot, we evaluated a spectrum of frameworks like LangChain, LlamaIndex, and others for chaining models and tools together and utilizing built-in connectors to vector databases and memory stores. Ultimately we chose to standardize on Vercel's AI SDK for its ease of use and seamless integration with our Next.js and serverless stack. Today we're using it in production and are satisfied with the choice.
To understand how LLM-driven apps manage continuity, it's important to recognize that human / LLM interactions are stateless by default (i.e. without an automatic context management mechanism). Using Vercel's AI SDK (or similar framework), each request is assembled by serializing the system prompt, user input, and previous message data, containing the full context for the next LLM response. With no default memory for the LLM to draw on between turns, it is up to developers to decide how to continually compress context to stay within token limits. It is a new style of development with a different request lifecycle - contrasted with a typical frontend app where you'd make a fetch call to an API and rely on the backend to manage session state, in an LLM-driven chat app, API interactions orchestrated by an LLM will have response data included in the context window.
This reality has led to the emergence of new engineering sub-disciplines focused on context shaping (preserving user sentiment, intent, and history under tight token budgets). Techniques like summarization, reference compression, and message prioritization are becoming baseline requirements, even for simple apps.

Reference Tables
One of our most efficient data saving techniques is extremely simple - providing query references to the client so it can fetch additional data on demand. For example, in our chat UI agents may render React components as part of a tool call result, with one of the most common use cases being a table of data. Instead of loading full API payloads into the context window, we provide some sample table data to the client for initial server-side rendering and a reference to the rest. This helps to keep large tool result payloads out of the context window and enabes AI-generated React components to interact directly with our backend API.
A transaction group can contain thousands of records. Instead of sending all of them through the model, we send only a lightweight reference and a small data sample:
- Tool call output: group metadata + the first N transaction IDs.
- Initial render: server uses those IDs to fetch details and render the first rows.
- Client behavior: as the user scrolls, React Query requests more IDs from the backend and hydrates them into full transactions.
Our custom hook implements this "reference table" pattern for use by AI-generated React components, providing infinite scroll functionality without bloating the model context.
// useReferenceTransactions.ts
import { useCallback } from 'react'
import { useInfiniteQuery } from '@tanstack/react-query'
import { getGroupDetails, getTransactions } from '../api/transactions'
export function useReferenceTransactions(groupId: string, initialTransactionIds: string[], pageSize = 20) {
const fetchPage = useCallback(async ({ pageParam = 0 }) => {
// PAGE 0: Use IDs provided in the tool call result (already in chat payload).
// This means no network request for metadata — acts like server-rendering a few initial transactions.
if (pageParam === 0 && initialTransactionIds.length) {
return getTransactions({ transactionIds: initialTransactionIds, page: 0, pageSize })
}
// SUBSEQUENT PAGES: Ask backend for more IDs by group reference.
// This keeps context window light: LLM only saw IDs, client resolves details now.
const group = await getGroupDetails({ groupId, page: pageParam, pageSize })
// Resolve those IDs into full transaction objects.
return getTransactions({ transactionIds: group.group.transactions?.ids || [], page: 0, pageSize })
}, [groupId, initialTransactionIds, pageSize])
return useInfiniteQuery({
// QUERY SHAPE: ['referenceTx', groupId]
// Each unique groupId gets its own infinite query.
queryKey: ['referenceTx', groupId],
queryFn: fetchPage,
initialPageParam: 0
})
}Reference tables like the one below (transactions) might have a small handful of initial, server rendered transactions but the majority are fetched via an Infinite Query. This pattern alleviates model context bloat, while giving the user a rich, scrollable interface inside the chat.

Keeping large volumes of data out of the context window is also a key benefit of MCP, which we are exploring for internal use. However, we do not think it will completely solve all of our problems, as we will also need to stay under context window budget by pruning and structuring data via relevance filtering, summarization, segmentation, structured injection, and strict budget checks, just at a different layer of the stack. Depsite critcisms of MCP, it may still have a place in our stack, but we don't believe it will be our definitive context window management solution.
Conclusion
Context window management is an inherent challenge and limitation when building products with LLMs. Even simple apps must be mindful of this fundamental constraint to prevent app crashes from token overflow and provide the right context at the right time to the LLM to best serve the user's needs.
Our Accounting Specialist Agent is now live in production helping our customers close their books faster every month. We've been continuously learning and refining our approaches as new context window crunching practices emerge in the field today. This is an exciting and challenging new development area, requiring product-focused engineering solutions to best facilitate customer objectives. If these problems interest you, discuss with us on X and reach out if you're interested in working with us!
Related

Entendre introduces a new AI Agent Marketplace to help businesses automate complex financial operations with customizable, specialized AI agents.

Entendre receives SOC2 Type 2 compliance

Entendre and Credit Cooperative partner to automate end-to-end borrowing and lending on-chain, while simplifying complex blockchain transactions.

Entendre has been acquired by MoonPay. We continue as-is for all existing and new customers, now with more resources, faster deployments, and deeper integrations.