ArchitectureA supervised JVM-class runtime — OLTP on seven engines, OLAP on three. AI-native, MCP-native, observable as plain SQL.Read the architecture
Está viendo la edición Perú. Está viendo la edición Colombia. You're viewing the Pakistan edition. Cambiar a la edición global →Cambiar a la edición global →Switch to the global edition →

Token-aware AI calls: accurate counts, bounded memory and one retry policy

Conversation memory is now bounded by tokens, not messages; token counts are estimated per provider and can be confirmed before a request is sent; and every provider call shares one retry policy.

An agent conversation fails in unglamorous ways: a single tool result carrying a large result set overflows the context window, a quota error is retried too fast to clear, a rejected request is retried three times to be rejected three more. This release addresses each of those at the platform level, so no feature built on it has to solve them again.

Counting tokens properly

  • Per-provider estimates. Token counts for Anthropic and Vertex AI models are estimated locally with a per-model-family correction, replacing a generic count that ran well short on prose and further short on the JSON and code that dominate agent conversations.
  • Exact when it matters. Where a provider offers it, the platform can ask for the exact cost of a whole request before sending it, confirming a conversation fits the context window before paying to find out it does not.
  • Every model says how it counts. Chat and embedding models both declare their tokenizer, so a budget is measured in the unit the provider enforces.

Memory bounded by size

  • Tokens, not messages. Conversation memory is capped by token count: the oldest exchanges are evicted first, then the largest remaining item is truncated, and the conversation always opens on a user turn.
  • Tool calls kept whole. A tool call and its results are evicted together, so the model never sees an answer without its question.

One retry policy

  • Consistent across providers. Every provider call shares one retry policy: a malformed or unauthorised request is not retried, while rate limits and transient failures are.
  • Waits taken from the failure. When a provider says how long to wait, the retry uses that figure instead of a fixed backoff that is too short for a per-minute quota.

These changes sit below every AI feature on the platform, so conversations, agents and tools inherit them without configuration.

See the feature →

← All posts