Building Parloa Zones: Repeatable regional deployments with central control
:format(webp))
Back in June, Parloa’s Senior Director of Engineering, Ítalo Vietro, wrote about deployment stamps at Parloa. He describes a stamp as "customer infrastructure in a box": an isolated unit that bundles everything a customer or group of customers needs, from compute and AI capacity to networking and monitoring, so what happens on one stamp stays on that stamp. Under the hood, each stamp is an Azure subscription, an AKS cluster and all the services in Parloa's runtime layer.
A few weeks ago, we launched Parloa Zones, our product offering built on top of deployment stamps. Zones lets customers deploy Parloa in the supported region of their choice and prove data residency, meaning transcripts, personally identifiable information (PII), analytics, plus the services that process them, stay inside that regional boundary, no matter where the agent is built.
Stamps gave us the foundation, but turning them into a product meant tackling three challenges: making new regions repeatable, keeping data regional while management stays central, and making sure each regional runtime keeps working even if it loses contact with the rest of the platform.
Where we started
Before Zones, Parloa had two regional stamps: app.parloa.com, hosted in Germany, and app.parloa.us, hosted in the United States. Each had a staging version for testing and validation, making four environments in total. Every service we deployed had a separate set of configuration files for each environment.
Keeping four sets of files was manageable. Creating new stamps was possible too, but it took coordinated efforts across engineering teams to provision infrastructure, deploy services, and test them.
With Zones, we wanted the ability to spin up a Parloa stamp in any supported region and maintain many more of them. Expanding to a new region couldn’t depend on asking a bunch of teams to generate new config. files.
Challenge 1: Making regional infrastructure repeatable
Parloa's tech stack is cloud-native. Almost all of our services run in Kubernetes, and we deploy them with Argo CD following the GitOps model. Both are highly stable projects maintained by the Cloud Native Computing Foundation. We run on Microsoft Azure and manage those resources with Terraform, following infrastructure-as-code best practices. We have been largely happy with those choices.
But take all the services we've built as part of the product, add the platform-level services we run (monitoring, autoscaling, networking, security scanning, and so on), and you get close to 100 services in a single deployment stamp. Setting all of that up by hand for every new region, with each team writing its own config files, wasn't realistic at the scale Zones needed. So we added a new layer of tooling to make it easy to "stamp out" a new infrastructure footprint on demand.
The stamp manifest
The core of that new layer of tooling is the stamp manifest: a YAML file that declares a stamp's key attributes, like which region it's in, whether it's a production cluster, and a t-shirt size for how big it is. A slightly simplified version looks like this:
apiVersion: "parloa.com/v1"
kind: "StampManifest"
metadata:
name: "production-canada"
owner: "tech-platform"
spec:
location: "canadacentral"
capacity: "medium"
isProduction: trueAdding one of these files to our stamps-catalog repository triggers a series of automations. They create the Azure subscription and the AKS cluster, configure networking and security settings, and deploy Argo CD to the cluster, so it can fetch and deploy every other service.
We do still need some per-region configuration, but we as the Tech Platform team can handle that without requiring all of the other teams to participate.
Crossplane for datastores
We also chose Crossplane, an open source framework that lets you declare and manage resources like databases, as Kubernetes resources. If a dev team runs a service that needs a Postgres database or a Redis cache, they can specify that in a few lines of YAML without touching Terraform:
resources:
caches:
- alias: config_cache
size: s
postgres:
- alias: knowledge
size: m
readers: trueOur Crossplane-based tooling provisions the datastores, sets up roles and access policies, generates credentials, and makes them available to the containers running in the cluster. Whether we have one stamp or several dozen, the team that owns the service doesn't need to do any extra work.
Designing for the teams who use it
The Tech Platform team worked closely with other engineering teams to understand what they needed. The goal was to keep things as simple as possible so teams weren't overwhelmed with detail, but still had the options and parameters necessary to build and deploy the services behind the AI agents our customers rely on. That was an iterative process with lots of feedback and adjustments.
What repeatable regional infrastructure enables: Adding a new region becomes a manifest file plus some focused work from one team instead of a cross-team infrastructure project.
Challenge 2: Keeping data regional while management stays central
As Ítalo described, Parloa's platform is built on two planes: a global control plane and a regional runtime plane. That split is what makes Zones work. If Zones is available in a dozen countries, that shouldn't mean a dozen different UIs are required to log into it. We want customers to be able to build and manage their agents in one central place, while those agents run and handle conversations in the region each customer chooses.
The control plane manages configuration: agent building, phone numbers, and channels. The runtime plane handles live activity: phone calls, chat sessions, and the conversation data they produce. Each regional runtime stamp receives its configuration from the control plane, then runs independently within its region.
To support Zones, we worked with engineering teams to define exactly which services belong in the control plane and which belong in each regional runtime. For each service, we reviewed all the data it stored and made sure runtime data, especially conversation transcripts, stays in the runtime stamp. Data analytics are also computed and stored in the stamp's region.
The split also changed how Studio, the interface where customers build and manage their agents, works. Some Studio API calls now go to the control plane, while others go to the runtime stamp in the customer's region. Studio looks up the right endpoint dynamically through an API.
What regional data and central management enable: Customers get a single place to manage every agent, and legal and security teams get a clear architectural boundary to point to when they need to prove where conversation data lives.
Challenge 3: Keeping regional runtimes running on their own
As part of the architecture, we wanted to make sure runtime stamps could run independently of the control plane because a caller on the phone with an AI agent can't be put on hold while a stamp waits to reach it. If a temporary network partition cuts a stamp off, the agent can still handle phone calls and chat sessions.
That meant building a caching layer into each runtime stamp, so all the agent configuration a conversation needs is already there the moment a call or chat starts. The caches update automatically in the background, rather than waiting for an agent to handle a call. If we didn’t take this approach, and a stamp only fetched agent configuration when a call came in, we’d run the risk of being unable to handle the call, if the connection to the control plane was down or slow.
With the caching layer, the stamp pings the control plane every few seconds, asking if there’s any new data to fetch. That’s configurable, with automatic backoff, to protect ourselves from accidentally overloading the control plane. In Studio, customers can see at a glance whether a stamp is running the latest agent configuration. It only shows as up to date once the stamp has fetched and validated it.
What independent runtimes enable: A network blip between the control plane and a region doesn't turn into a dropped customer call.
What it takes to add a new region
So we've automated the tooling. Spinning up a stamp in any country in the world should be as easy as snapping our fingers, right?
Unfortunately, it's not quite that simple. First, we run on Microsoft Azure, so we're limited to the regions Azure supports. Even if we wanted to create a stamp in the Maldives (and schedule a team offsite there to check on it), we can't, at least until Microsoft opens a datacenter there. Alas, no white sand beaches for us.
Even when a country does show up on the supported list, our work isn't done. We have to verify that the region has every resource we need. Hyperscalers roll out new features and capacity to regions at different rates, so we cross-check every cloud resource we use, confirm it's supported in that region, and check with Azure on available capacity.
That holds especially true for LLMs. Models hosted by OpenAI or Azure may not be available in every region. Or they may be available without enough provisioned throughput units (PTUs), the reserved capacity that guarantees a consistent level of service. When you're talking to an agent on the phone, you don't want delays or gaps in the conversation because there isn't enough compute behind it.
It's our job to do the research and due diligence, so every customer gets a high level of performance and availability. Once we’ve verified a region will work, getting it up and running should be easy and automated.
Going global
Today, in regions with the Azure services and model capacity we need, standing up a new stamp is now largely automated. Getting there took a manifest-driven way to stamp out new regions, a clean line between where agents are managed and where conversations happen, and runtimes that keep working on their own.
And if Microsoft ever opens a data center in the Maldives, we’ll be ready.
:format(webp))
:format(webp))