
Sammy Emir
"Dream big, start small, but never stop moving forward."
Inside a plasma etch chamber, a wafer worth more than a car is exposed to corrosive gas, extreme heat, and a plasma environment where energetic ions and reactive chemical species remove material with nonometer-scale precision. Surrounding that wafer is a set of components engineered to absorb the damage on its behalf. They are called process kits and their job is to shape the plasma, protect critical chamber hardware and maintain process performance while withstanding gradual wear over time. The industry they support is now enormous. Global semiconductor sales are projected to reach $1.51 trillion in 2026, with memory alone climbing past $800 billion on the back of AI demand. Almost none of that output is possible without the unglamorous hardware sitting a few millimeters from the wafer edge, and when that hardware fails, the chips stop.
Sahiti Nallagonda has spent more than 10 years designing it. A senior mechanical engineer and inventor with 5 patents, she has served as the lead mechanical engineer and technical owner for process kit design, qualification, and sustaining programs across multiple etch platforms and product generations, working the boundary where materials science meets high-volume manufacturing. Her hardware is deployed in more than 1,000 chambers and is used daily by engineering, lab operations, service, and supplier teams across several organizations.
We spoke with Sahiti about why the parts nobody notices decide whether a fab hits its numbers, what happens when a component erodes in the wrong place, and how she designed hardware that replaced a decades-old approach to controlling the wafer edge.
Most people picture chipmaking as lithography and cleanrooms. What are they not seeing?
They are not seeing the consumables. Inside an etch chamber there is a set of parts surrounding the wafer, rings, liners, shields, and their entire job is to shape the process and take the abuse. The plasma is aggressive by design because it removes material with atomic-scale precision. But it does not politely stop at the wafer. It attacks everything it touches, including the hardware holding the wafer in place.
So those parts erode. That is expected and it is the point. What is not expected is for them to erode unevenly, or shed a particle, or shift by a fraction of a millimeter. Deviations as small as a few tens of hundreds of microns can alter local plasma conditions and affect wafer-edge performance. You end up with a die that fails at the perimeter of a wafer that cost thousands of dollars. A part that is technically a consumable turns out to be a yield-limiting component. Most people outside the industry have never heard of it. Everyone in the industry has an opinion about it.
You led process kit development across several platforms. What problem were you actually solving?
Several, and they compete with each other. Particle accumulation was one. Process byproducts builds up in the chamber over time, and if it accumulates in the wrong geometry it eventually flakes off and lands on a wafer. We addressed that with a wide-gap ring design that changed where deposits could collect and how they could be cleaned out. Erosion was another. A part made from a single material erodes at a single rate, which is fine until the erosion profile starts distorting the process. We moved to hybrid kits that combine quartz and silicon carbide, placing the more plasma resistant material in regions experiencing the highest ion and chemical loading.
Then there is the part I am proudest of. On some plasma etch architectures, wafer-edge behavior is influenced through capacitive coupling between chamber components. I worked on hardware that enabled direct electrical biasing of a process ring, providing additional control over wafer-edge conditions. I worked on hardware enabling pulsed voltage technology. It was first-of-a-kind hardware, and it gives process engineers real tunability at the wafer edge instead of a fixed condition they have to design around. That demonstrated how process-kit hardware can evolve from purely passive consumables into electrically integrated elements that contribute directly to process control.
Walk me through what makes this mechanically difficult.
Everything is fighting everything. The part has to survive high vacuum, Elevated temperatures and thermal gradients that can vary significantly across chamber components during manufacturing and maintenance cycles. It has to hold tolerance while it is doing that, because if it moves thermally in an unplanned direction, the gap changes and the process changes. It has to be installable by a technician in a fab at 2 a.m. without special tooling, and it has to go in the same way every time, because installation variability shows up as process variability. And now, with the wired ring, It has to maintain a stable electrical path across interfaces that are exposed to thermal cycling, contamination and plasma byproducts.
So the design space closes fast. Pick a material that survives the chemistry and it may not survive the thermal cycling. Solve for thermal and you may create a tolerance stack that no supplier can hold at volume. My job was rarely finding the clever solution. It was finding the solution that was still viable after you accounted for the supplier, the technician, the thermal budget, and 5 years of erosion. Most designs that look elegant on a screen die on one of those.
What went wrong that you did not see coming?
The failures that hurt are the ones that appear at scale, not in the lab. A kit qualifies beautifully on a test chamber, then goes into a fab and installation variability across sites produces results nobody can explain. Or a tolerance that was fine on the first article stacks badly against a mating part at volume. In a fab, that is expensive in a way people underestimate. Maintenance discipline drives overall equipment effectiveness directly, and a healthy fab wants a maintenance ratio of 4.0 or higher, meaning scheduled downtime should outweigh unscheduled downtime by at least 4 to 1. Every unplanned chamber event drags that ratio the wrong way.
That is where a lot of my work went, and it is the part that does not photograph well. Better fixture designs. Better calibration methods. Making installation repeatable so the same part behaves the same way in Taiwan and in Arizona. Across several programs, resolving mechanical, tolerance, and installation issues prevented roughly 50 chamber rework and field retrofit events. Each event can cost $50,000 to $70,000 in labor, parts, and lost lab time, and that is before you count the wafers that did not get processed. Nobody gives you an award for the failure that never happened. It is still the most valuable work I do.
You also review academic research in manufacturing. What do you look for in a paper?
Whether the authors have ever been in a fab. I peer review for the Journal of Intelligent Manufacturing, including a comparative study of heuristic and learning-based methods for anomaly detection in overhead hoist transportation systems, the automated tracks that move wafer carriers around a fab. It is a good example of the gap I look for. You can build a model that detects anomalies with impressive accuracy on a clean dataset. Whether it detects them early enough, and with few enough false alarms, that a fab would actually let it interrupt production is an entirely different question.
So I push on the constraints. What happens when the sensor drifts. What does a false positive cost the fab. Would anyone act on this output at 3 a.m. The strongest papers are honest that a method has limits, and they characterize those limits precisely. The weaker ones report an accuracy number and stop. Manufacturing is unforgiving in a way that a benchmark dataset is not, and reviewing keeps me sharp about the difference. It also keeps me current on approaches I might otherwise never encounter from inside a hardware group.
Hardware engineering has a reputation for being slow and incremental next to software. Fair?
Partly fair, and I think the framing is wrong. Hardware is slow because physics does not accept a patch. If a material erodes, you cannot push a fix over the air. You redesign, you re-qualify, you re-release, and the cycle takes months because it must. That is not conservatism, that is the cost of building something that has to hold up under conditions that would destroy almost anything else.
But incremental is the wrong word. Replacing capacitive coupling with a direct electrical connection at the wafer edge is not incremental. It changes what process engineers are allowed to do. The pace of hardware innovation is set by the qualification cycle, not by the ambition of the idea. What I would say to anyone dismissing this field is simple: every AI accelerator and every high-bandwidth memory die on the market passed through a chamber whose behavior was determined by a set of mechanical parts somebody had to design. The software runs on the hardware. The hardware runs on the parts.
Where does this work go next?
Toward more active hardware. Investment is running far ahead of where it was, with 300mm fab equipment spending rising 18% to $133 billion in 2026 and $151 billion in 2027, and the demands on the chamber are rising with it. Advanced transistor architectures and taller memory stacks need tighter edge control and less variation than the current generation of process kits can deliver. The passive part model, where a component just endures the plasma, is running out of headroom.
So the direction I am working toward is process kits that participate: electrically integrated, using new material combinations and contact strategies, tunable rather than fixed. The wired ring was a first step, not a destination. What I want, specifically, is for a process engineer to be able to adjust conditions at the wafer edge as deliberately as they adjust anything else in the chamber, instead of accepting whatever the hardware happens to impose. That is a mechanical design problem, and it is a long way from solved.
Most of the data a company creates is never meant to last. Session tokens, event logs, sensor readings, cache entries, the temporary exhaust of running software. It arrives in enormous volume, stays useful for minutes or hours, and then becomes clutter that costs money to store and slows everything down. The world generated roughly 181 zettabytes of data in 2025, and a substantial share of it consisted of short-lived operational data that was only useful for minutes, hours, or days. Handling that gracefully, letting data expire on its own without an engineer babysitting it, turns out to be one of the quietly hard problems in modern databases.
Varsha Ganesh has spent 14 years building the systems that make this look easy. She is a Senior Software Development Engineer who works on large-scale distributed database infrastructure and cloud services. One of the most significant projects she led was the implementation of Time-to-Live (TTL), the capability that allows data to expire automatically, for a managed distributed database platform. The effort required coordinating four engineering teams while redesigning core storage-engine components to support expiration at production scale.
We spoke with Varsha about why automatic data expiration is harder than it sounds, what it takes to delete records at massive scale, and how a single missing feature can block an entire class of customers from moving to a new system.
People think of databases as systems for keeping data. You spend a lot of time on the opposite problem. Why is deleting data so hard?
Because deletion at scale is a promise you have to keep forever. Storing a record is easy. You write it and move on. Telling a customer that a record will disappear at a specific time, reliably, across a distributed system that never stops taking traffic, is a much heavier commitment. You have to track expiration for every row, find the expired ones without scanning everything, and remove them without disturbing the live workload sitting right next to it.
The other reason is that data people intend to throw away still shows up in huge quantities. A team running session state or telemetry does not want to write cleanup jobs, monitor them, and debug them at 2 a.m. when the cleanup falls behind. They want the database to handle it. When it does not, that work does not disappear. It just moves onto an engineer's plate, and it stays there.
Many customers were waiting for native Time-to-Live support before migrating. Why was that one feature such a blocker?
The platform was a fully managed distributed NoSQL database used by customers migrating from self-managed environments. One of the most requested capabilities was native Time-to-Live support. Many customer applications were already designed around automatic expiration. They stored session information, telemetry, temporary events, and other short-lived records that were expected to disappear without scheduled cleanup jobs.
Without native expiration, those customers had to build and maintain their own cleanup processes, which became a major obstacle to migration. So although the feature itself was TTL, the real objective was removing one of the biggest adoption barriers. Once that capability existed, organizations could move existing workloads without redesigning how their applications handled short-lived data.
Walk me through what makes expiration at this scale technically hard.
Start with the volume. IoT devices alone were on track to produce around 90 zettabytes of data a year, and that kind of data, telemetry and events and readings, is exactly the kind that carries an expiration date. In the largest production environments, the platform processes millions of requests every second, and every write now had to account for expiration metadata, even for records that never expire. That meant the work lived directly in the core write path that every request passed through. There was no isolated subsystem where the feature could be hidden. That value has to be set explicitly on all of them. So this work lived in the core write path that every single request goes through. There was no edge of the system to tuck it into.
Then there is the deletion itself. You cannot find expired data by scanning the whole dataset, that would be ruinously expensive. Physical cleanup happens as part of the storage engine's existing compaction work when it reorganizes and compacts data on disk. We rearchitected the storage and compaction layers so that expired records get cleaned up during that process, and we had to handle the conflict resolution carefully, because a record being written, updated, and expired can race in ways that produce the wrong answer if you are not deliberate about ordering.
Changing the core write path of a live system processing millions of requests per second sounds terrifying. What nearly broke?
The scariest part was that there was no clean slate. This was a live production service with customers already depending on it, and I was proposing to change how every record was stored. Get that wrong, and you don't break one feature. You break all of them. So the work had to be introduced beneath the existing architecture, preserving the original behavior while the new functionality was phased in, completely invisible to customers using the service.
The other hard part was coordination. Time-to-Live touched the storage engine, compaction, the query layer, and the background cleanup processes, which meant four engineering teams had to move in lockstep. When you're the one holding the design together, the technical decisions are only half the job. The other half is making sure multiple teams with different priorities agree on the same sequencing and the same tradeoffs. Getting people aligned was every bit as difficult as getting the code right, and it mattered just as much.
How has judging emerging AI projects influenced the way you approach your own engineering work?
It keeps me honest about what actually matters. As a judge for the Builders of Tomorrow AI Super Hackathon, I watch teams build under real-time pressure, and the pattern that separates the strong teams is clarity about the problem, more than technical firepower. The teams that win can tell you precisely who they are helping and what they are removing from that person's day.
That is the same instinct that made TTL worth doing. The feature was not glamorous. But it was the exact thing blocking real customers from a real decision. Judging reminds me that engineering is at its best when it is pointed at a specific, well-understood problem, not when it is showing off. Most of the impressive work I have seen, in a hackathon or in production, is impressive because it is precise, not because it is complicated.
Where is the managed database space heading, and what still holds customers back?
Adoption is still driven by parity, not novelty. The NoSQL database market is on track to reach $69.09 billion by 2031, and most of that growth is organizations modernizing legacy database environments onto managed ones. But they only move when the managed option does everything the old one did. A single missing capability, TTL, a particular consistency mode, a specific query pattern, is enough to stop a migration cold. Customers judge you by your worst gap, not your best feature.
So the real work in this space is unglamorous completeness. It is closing the small holes that individually look minor, and collectively decide whether someone can trust your system with their production workload. The interesting part is rarely a capability nobody has seen before. It is making the boring, expected things work flawlessly at a scale where flaws are expensive.
What's the part of this problem you still think about?
Data lifecycle is going to get more demanding, not less. As the volume of short-lived data keeps climbing, the systems that store it have to get better at forgetting on purpose, and forgetting is genuinely harder than remembering. There is real engineering left in making expiration cheaper and more precise, so that a customer can trust the database to clean up after itself at any scale without a second thought.
What I keep coming back to is that the best infrastructure is the kind nobody notices. When TTL works, no one thinks about it. Data appears when it should and is gone when it should be, and an engineer somewhere got to sleep through the night instead of babysitting a cleanup job. That is the goal I care about: building something so reliable it becomes invisible.
1 Like
Comment

An AI agent can draft a contract, reconcile an invoice, and move money across three countries before lunch. What it cannot reliably do is know that a postal code in Japan vs Jordan follows a different logic than one in Sao Paulo, or that a tax identifier valid in Berlin is malformed in Mumbai. This matters when organizations are scaling AI products across global markets, because writing new validation code for every market is not sustainable. The data those agents touch, including addresses, tax identifiers, names, and currency, changes from one jurisdiction to the next, and the rules governing it are not suggestions. They carry the force of law.
Sumeet Ram has spent more than 8 years in the technology industry, building systems across insurance, low-code automation, international lending, and global payments. A Senior Technical Product Manager at one of the world's largest digital payments platforms, she leads internationalization work that spans more than 200 markets, the layer most users never notice until a name is formatted incorrectly or an address field rejects a valid home. For example, in many Asian countries, the family name comes first, whereas in many Western countries, it comes last. Her current focus is what happens when AI agents begin touching that layer directly. The answer, she argues, decides whether automation becomes a compliance asset or a compliance liability.
The Problem With Letting Agents Reason
Large language models do not retrieve facts. They predict text, which means they approximate, and in regulated work approximation carries a measurable failure rate. Models hallucinate on as many as 41% of finance-related queries. A chatbot that invents a plausible transaction or cites a regulation that does not exist is not a rare edge case. It is the predictable behavior of a system asked to guess.
Ram saw the pattern early. When agents are handed locale rules, such as how a Brazilian tax identifier is structured, which characters a Japanese address allows, or where a decimal separator belongs, they often rely on patterns learned during training rather than the deterministic rules defined by an organization's systems. While AI can retrieve and reason over vast amounts of information and may even suggest the correct answer, applying that knowledge consistently within enterprise systems is far more challenging. The output looks confident and is sometimes wrong, and in payments a wrong address or a malformed tax ID does not fail quietly. It surfaces downstream as a rejected transaction, a compliance flag, or a customer locked out of a product that was built for them. Her position was blunt: locale logic is deterministic, and deterministic rules should never be left to a probabilistic model.
"Locale rules are not the kind of thing you want a model to have an opinion about," says Sumeet Ram. "There is a correct answer for how an address works in each market, and the job is to make the agent use it, not approximate it."
Why Close Enough Breaks at the Border
59% of online shoppers now buy from retailers outside their home country, and that demand only converts when the local details are right. A name field that cannot hold the customer's name, currency that displays with the wrong separator, a tax ID in the wrong format: each one reads to the user as a product that was not built for them. The failure is rarely dramatic. It is a rejected form, a flagged record, or a checkout that quietly does not complete.
This is the problem Ram has been working on since joining PayPal over the last 18 months, well before AI agents entered the picture. The internationalization platform she supports handles currency formatting, date and time conventions, address verification, name validation, locale metadata, language and regional preferences, character encoding, number formatting, and many other capabilities across more than 200 markets. Each of those capabilities exists because products are often designed with assumptions from the market where they were originally built, but those assumptions do not hold everywhere. A product designed with a US-first mindset, for example, cannot deliver a native experience in India, just as a platform originally built for Sweden or China must evolve to feel local in every market it serves. The validated functions behind that platform encode years of market-specific rules. The question she set out to answer was how to put that same validated logic in front of an AI agent without letting the agent paraphrase it.
"Globalization gets treated as a chore, something you bolt on at the end," Ram explains. "I have always thought of it as infrastructure. If it is infrastructure, an agent should be able to call it the same way it calls anything else it trusts."
Turning Locale Logic Into Something Agents Can Call
The Model Context Protocol, released as an open standard in late 2024, gave the industry a shared way for AI agents to use external tools instead of reasoning their way through every task. Adoption was fast, and it has since become the default way agentic systems connect to the software around them. The protocol is usually described as plumbing, a cleaner path between an agent and a tool. Its more useful role is as a boundary.
Ram led the work to build her company's first internationalization server on the protocol, scoping it from the standard's earliest public months to a production launch roughly 9 months later with a team of 5 engineers. The design choice that defined the project was narrow and consequential: every tool the server exposes is anchored to an existing, production-validated function. An agent calling the server to validate an address or format a currency does not receive a hint it then completes on its own. It receives the verified answer. The agent stays in the execution path and is kept out of the reasoning path on data that regulators care about.
"The protocol is the integration story everyone tells," Ram notes. "The part I care about is that it can also be a constraint. You can use it to guarantee an agent never improvises on something that has a single right answer."
The Decisions That Keep It Honest
Most agentic projects do not fail on the model. They fail on governance, on unclear value, and on the absence of controls that decide what an agent is and is not allowed to do. The hard part is rarely getting an agent to act. It is drawing the line between the work it can own and the work it must never improvise.
That line is where Ram spent much of her effort. Early proposals for the server included tools that returned partial answers and let the model finish them, and she rejected them, because a half-answered locale rule is the exact place a hallucination hides. Capabilities were prioritized by what an agent actually needs rather than by how much of the underlying system could be exposed, which kept the launch shippable. Discoverability, onboarding documentation, and review with security and compliance partners were scoped as part of the product, because a tool nobody can find or trust does nothing to reduce risk.
"The instinct is to let the model help out where the tool is incomplete," Ram observes. "That instinct is exactly what you have to design against. Either the answer is validated or it is a guess, and on compliance data a guess is a defect."
A Pattern That Reaches Past Addresses
The same logic applies wherever AI meets regulated data. Identity verification, know-your-customer checks, payments, and risk all rest on rules that are jurisdiction-specific, auditable, and unforgiving of error. They are the surfaces where letting an agent approximate would be most tempting, and most dangerous. They are also where a validated tool earns its place.
Ram describes the internationalization server as a template rather than a one-off. The pattern, exposing deterministic validated logic as tools an agent consumes rather than interprets, extends directly to identity, payments, and risk. Inside her company it has already become the reference for how other compliance-sensitive systems should meet AI agents, and the next phase of her work connects these servers to one another so an agent can compose validated capabilities without ever leaving the verified path. The goal is not a smarter agent. It is an agent that knows the difference between a fact it can look up and a fact it should never invent.
"The internet solved connectivity. It did not solve comprehension," Ram reflects. "If we want agents to work for someone in Sao Paulo as well as they do for someone in San Francisco, we have to hand them the rules, not ask them to guess. That is the whole job."
1 Like
Comment
Every technician a telecom carrier sends into the field costs money before any actual work starts, because somebody has to drive there. One route is nothing. Tens of thousands of them a day, across a national network, and small mistakes about who goes where pile up into millions of wasted miles and a lot of people sitting home waiting on a service window. An entire category of software now exists to solve this problem, and it isn't small. The field service management market is expected to grow from $5.10 billion in 2025 to $9.17 billion by 2030, at 12.5% a year, largely because doing it by hand stopped working a while back. What's still an open question is how much of the actual decision you hand to a machine.
Mundakkapatta Dileep Kainary has been at this problem for more than 22 years. He's an Application Architect at IBM and a Senior Member of the IEEE. Much of his career has focused on the systems that decide how large telecom operations move people and work across a network. One of those systems became the focus of our conversation: an AI platform he led that changed how one of the biggest field forces in U.S. telecom picks which technician to send where.
We talked with Dileep about what it takes to teach a field operation to make that call in real time, the parts that nearly fell over, and how far he thinks the automation should actually go.
Field dispatch sounds like a scheduling problem. Why is it harder than it looks?
On paper, sure, you're matching technicians to jobs. In practice it's a decision problem, and the inputs never sit still. Where the tech is this minute, what the job actually needs, traffic, what parts are on the truck, whether the last job ran long. A dispatcher can juggle maybe a dozen of those at once and do fine. A national carrier throws far more at you than any one person can keep in their head.
So the whole thing quietly goes sideways. Everyone optimizes for what's in front of them, sends the closest person, and the system as a whole gets worse without anybody deciding it should. The nearest tech wasn't the right skill match, so the job bounces. Two trucks cross the same neighborhood because nobody could see the full board. It never looks like a failure. It looks like miles and overtime, and most operations just file that under the cost of doing business. I think that's a mistake.
You led an AI platform built to solve exactly that. What did it actually do?
The dispatch optimization platform looked at the whole board at once. It pulled live and historical signals together: skills and certifications, real-time location and traffic, fuel, customer history, what the day's workload actually looked like. Instead of asking who's closest, it asked what single assignment, given everything happening right now, ends up best for the customer and for the network. I led the architecture and delivery, working with the dispatch teams, the data scientists, and the business side to turn all that operational mess into decisions the system could make itself.
And the numbers moved. Technician travel dropped by more than 20%. Over the life of the program that came to something like 51 million pounds of CO2 that never went into the air, which was ultimately a consequence of driving fewer miles. Customers waited less, more jobs got fixed on the first visit. I'll be honest, the customer number is the one I care about, but it and the carbon number come out of the same decision, so you get both or you get neither.
Making that call in real time, across a live network, is a serious engineering problem. How did the system actually decide?
Underneath it was a distributed, event-driven setup. Signals came in continuously and the platform reworked assignments as things shifted, instead of on some fixed cycle. We built it as separate services on purpose, so the routing logic and the data ingestion and the models could each get swapped or updated without pulling the whole platform down. That wasn't a nice-to-have. The operation runs around the clock, so you're doing maintenance on a moving car.
What we did back then was decision support, a human still pressed the button. The industry's moving past that now. By the end of 2026, 40% of enterprise applications will include task-specific AI agents, up from under 5% a year earlier. Dispatch is honestly one of the better places to try it, because the decision has clear edges, you find out fast whether you were right, and a bad call shows up as an actual number. The model was never the scary part. Handing it the authority to act, that's where it got interesting, and slow.
What nearly broke it? Where does a system like this go wrong?
Data quality, mostly. An optimization engine is only as good as what you feed it, and field data is messy in ways you don't appreciate until you're standing in it. A status says a tech is free when they're still wrapping up. The traffic feed is ten minutes behind the actual road. A skill code claims someone can do a job they've never been trained on. Feed a confident model bad inputs and it makes a confident bad call, fast, again and again. We spent far more time early on cleaning and reconciling signals than we ever spent on the routing math.
The other thing that nearly sank it was trust, the human kind. These dispatchers had run the board on gut for years, and the moment a system starts telling them they're wrong, they push back, which is fair enough. If the tool overrides someone and won't show its reasoning, they stop using it, or they quietly route around it. So we made it explain itself. Here's the technician it chose. Here's what it gave up to make that choice. People really underrate that part. The algorithm was maybe a fifth of the work. The rest was getting people to trust it, and a lot of good systems die right there in the pilot because nobody planned for it.
You are also a named inventor on a U.S. patent, for an interceptor architecture used in cloud migration. That is a different kind of work from dispatch. What connects the two?
On the surface they look like two different jobs, but they come out of the same instinct. That patent covers an intelligent interceptor architecture, a way to connect modern cloud platforms to old legacy systems without touching the legacy side. You drop a smart layer in the middle and let it handle the translation. Dispatch optimization follows the same architectural pattern. You put a decision layer between messy real-world inputs and the action you take, instead of trying to clean up every input at its source. Most of my career has been some version of that one idea.
What doing both taught me is that the hard part is almost never the clever core. It's the edges. In the interceptor work it was dozens of legacy protocols that each behaved a little differently. In dispatch it was the dirty field data. The real engineering is making one clean design survive contact with a hundred messy special cases, and honestly that's the part I like most. A system that only works when the inputs are perfect isn't finished, it's a demo.
The emissions savings on that project were significant. Is sustainability a driver in this work, or a byproduct?
Byproduct, mostly, and I'd rather say that plainly than dress it up. Nobody green-lit the Dispatch Learning Engine to save the planet. They wanted lower cost and faster service. But efficient routing and lower emissions are the same thing with two different labels on it. Fewer miles is less fuel, that's the whole trick. Tune the operation for time and money and the carbon comes down whether you were aiming at it or not.
Where it gets interesting is the scale. Transportation is the single biggest source of greenhouse gas emissions in the US, about 28% of the total, and a good chunk of that is commercial fleets still making routing decisions that are far less optimized than modern software makes possible. If software can pull double digits out of the miles a national field force drives, and it can, then this is quietly one of the more useful climate levers around, mostly ignored because it never got sold as one.
Where does this go next? What is still unsolved?
The next piece is closing the loop. Right now most of these systems recommend and a person confirms. Where it's headed is systems that decide and act inside guardrails you set, then tell you what they did and why. That's a bigger jump than it sounds, because the question stops being can the model make a good call and turns into how much rope am I willing to give it, and how will I know when to take it back. That boundary is the part I actually want to work on. Chasing one more point of accuracy bores me next to it.
What nobody's really cracked is the weird middle, the cases no training data ever saw. A storm takes out half a city and every route you planned is useless. A crew calls in sick. Some customer situation that fits none of your categories and just needs a person who can improvise. People are still better at those, and any system that pretends otherwise looks brilliant right up until the day it doesn't. So what I'm putting my time into is operations where the machine runs the routine at full speed and kicks the genuinely strange stuff to a human, cleanly, with everything they need to decide. Get that handoff right and the whole automation-versus-judgment fight mostly dissolves. That's the harder version of the problem, and the more honest one.
1 Like
Comment
Technology modernization in the public sector is no longer a matter of “if”—it’s a matter of “how.” As federal systems face rising demands for speed, security, and scalability, the traditional model of large-scale tech replacement is giving way to something more pragmatic: intelligent, phased transformation. At the forefront of this shift is Chandra Sekhar Kondaveeti, a technology expert, whose work in modernizing mission-critical platforms like the U.S. Department of Labor’s workers' compensation system is setting a new benchmark for scalable, secure, and citizen-first solutions.
With over two decades of experience in enterprise architecture, cybersecurity, and government tech leadership, Chandra has helped agencies move away from monolithic systems and into agile, interoperable environments—without compromising operational continuity.
“Scalability isn’t just about performance—it’s about resilience. It’s the ability to meet increasing demand without rewriting everything from scratch,” Chandra, who was also a Session Chair at the 2024 IEEE International Conference on Augmented Reality, Intelligent System, and Industrial Automation, notes.
Engineering Modernization Without Disruption
In the world of government IT, reliability isn’t a luxury—it’s a mandate. Agencies like the Department of Labor’s Office of Workers’ Compensation Programs (OWCP), which processes more than 100,000 claims annually, must serve thousands of users each day, including federal employees, healthcare providers, and case managers. In such a high-volume and compliance-driven environment, even the slightest system failure can lead to cascading delays in care and compensation delivery.
Rather than advocating for a costly system rebuild, Chandra took a more strategic approach. He led the transition to a microservices-based architecture, allowing the platform to evolve piece by piece while maintaining uninterrupted service. By replacing monolithic workflows with modular, containerized services, his team was able to modernize the system incrementally. Secure API layers replaced outdated communication channels, while real-time document tracking gave providers and administrators much-needed visibility into the status of medical submissions.
The results were immediate and measurable. Medical authorization processing times dropped by nearly 30 percent. Provider communication accuracy doubled, improving both care coordination and claim resolution. Perhaps most significantly, automation initiatives reduced manual form handling by 40 percent—freeing up critical time for both frontline workers and support staff.
“The goal is always to reduce friction—both for the system users and for the teams maintaining it,” says Chandra. “If it works behind the scenes without users noticing, it’s working well.”
Innovation Meets Security: Building Trust Through Tech
Modernization isn’t just about convenience—it’s also about trust. As digital systems become more integral to public services, the pressure to balance innovation with security is greater than ever. For Chandra, that balance starts with engineering discipline. In every system he touches, security isn’t an afterthought—it’s built into the foundation.
From OAuth2 protocols to automated vulnerability detection and fine-grained, role-based access controls, Chandra’s architecture choices ensure that even the most agile systems remain defensible. His proactive stance reflects a broader industry imperative: to build solutions that not only meet today’s compliance standards but evolve to address tomorrow’s threat landscape.
“You can’t innovate at the cost of security,” Chandra explains. “The most impactful solutions are the ones that are both forward-looking and inherently secure.”
This philosophy is reflected in his academic work as well. In his scholarly paper titled Advanced Performance Diagnostics in Modern Architectures: Thread Dump Analysis as a Key to Sustainable Scalability, Chandra explores how systems can be designed to maintain peak performance under pressure—without compromising on safety or scalability.
Human-Centered Transformation: Going Paperless with Purpose
One of the standout achievements from Chandra’s recent work was the successful implementation of a completely digital medical provider portal, enabling healthcare partners to submit, track, and manage claims without paper-based processes. But this wasn’t just a UI upgrade—it required restructuring legacy document workflows, integrating real-time status updates, and enabling seamless communication with existing claim systems.
The result? Tangible improvements across the board:
~50% decrease in claim-related disputes caused by document lag.
Drastic cuts in administrative overhead and mailing costs.
Enhanced provider satisfaction and engagement metrics.
This shift echoes wider trends. According to Gartner, by 2026, more than 70 % of government workflows will be digitally native. Chandra’s contributions place him—and the OWCP project—well ahead of that curve.
“Every digital change should create a better experience for someone—be it a citizen, a case worker, or a provider. If it doesn't serve people, it’s not transformation,” says Chandra.
Sustainable Transformation at Scale
As government agencies evolve to meet the demands of a digitally connected public, the success of modernization efforts will depend not on flashy tech—but on adaptable architectures, secure frameworks, and human-centered design. Chandra Sekhar Kondaveeti, whose work is also featured in MSN, emphasizes sustainable transformation: systems that evolve without disruption, prioritize user trust, and deliver measurable impact.
“Modernization isn’t one big leap. It’s a series of smart, strategic moves that add up to real impact,” he adds.
1 Like
Comment

Every serious AI project eventually runs into the same wall, and it is rarely the model. Research teams at universities and frontier labs can design a promising system in weeks, then lose months finding qualified people to label and judge its outputs, and more time still wiring up access to the handful of commercial models they want to compare it against. The market that supplies that labor shows the strain. The global data annotation tools market was valued at $1.69 billion in 2025 and is on track to reach $14.26 billion by 2034, growing close to 27% a year. The spending keeps climbing, yet the work behind it stays fragmented, costly, and hard to reproduce from one lab to the next.
Zihan Wang, a Forbes Technology Council Member, has spent his career on that overlooked layer between raw models and usable research. As co-founder and chief research officer of a Silicon Valley AI data infrastructure company, and the founder of an open research foundation that publishes evaluation benchmarks with academic partners, he works on the parts of the pipeline most teams treat as plumbing: where the data comes from and how competing models get compared on equal terms. His central claim is that the real bottleneck in AI research has moved off the algorithm and onto the infrastructure that feeds it.
The Human Layer Nobody Wants to Own
Modern AI systems still depend on human judgment at the points automated checks cannot reach: writing reference answers, grading model outputs, catching the mistakes a scoring script will happily wave through. That expertise is badly scattered. A lab studying medical reasoning needs annotators who understand clinical context; a team evaluating code agents needs working engineers, not generalists with a rubric. Most groups source these people ad hoc, through personal contacts or short-term contractors, which leaves quality uneven and nearly impossible to reproduce when the next project begins.
Wang's answer was to make that human layer a managed system instead of a scramble. He built a global network of vetted experts that research teams can draw on directly, covering everything from image and video to code, math, and specialized professional domains. People are sourced, screened, trained, and quality-checked before they touch a research task, with onboarding delivered through structured courses rather than a one-page brief. The network now spans more than 5,000 qualified experts, and it has become the layer that frontier labs and university programs at Stanford, MIT, and Harvard reach for when they need domain-qualified human judgment at scale.
“The hard part of evaluation was never the model. It is finding people who understand the task and will judge it the same way twice,” Zihan Wang says. “The moment you treat human expertise as infrastructure instead of a favor you call in, the whole process gets faster and a lot more honest.”
Fifty Models, Fifty Integrations
The model side of research carries its own tax. Comparing systems fairly means running the same test across many of them, but every provider ships a different interface, different limits, and different quirks, so one benchmark can burn weeks of engineering before it produces a single number. Concentration makes it worse. More than 90% of notable frontier AI models released in 2025 came from industry labs, each with its own access terms, which leaves academic researchers integrating and babysitting infrastructure they had no hand in building. For a small group without dedicated engineers, that overhead alone can decide which experiments ever get run.
To strip out that overhead, Wang built a unified access layer that puts more than 50 frontier models behind one standardized interface, including the major families from the largest labs. A researcher writes against a single API and routes the same prompt to dozens of models without rebuilding anything, while a normalization layer absorbs the formatting and behavior differences between providers. Setup that used to take weeks now takes hours. The same layer holds responses stable under load, which matters when an evaluation fires thousands of parallel calls and one flaky endpoint can quietly poison the results.
“Researchers should be testing ideas, not maintaining 10 different SDKs,” Wang explains. “When access stops being a project of its own, people finally run the comparisons they care about instead of the ones that were simply easier to set up.”
Proof That Holds Up in Public
Infrastructure earns its keep only through the work it lets other people reproduce. A benchmark built on inconsistent labeling or an unstable model pipeline yields numbers that look exact and signify almost nothing, because no outside team can reconstruct how they were reached. The signal that counts is published work, run on documented data under stated conditions, that survives scrutiny from people with no stake in the outcome. A strong score inside a private setup proves very little.
That standard runs through the research these platforms have supported. Wang co-authored a peer-reviewed benchmark that tests whether AI agents can carry what they learn across separate sessions and use it to make better decisions later, built on human-crafted tasks with university research collaborators. Its core finding is blunt: agents that score near the top on older memory tests collapse once they have to act on what they supposedly remembered, exposing a gap the field had been measuring around for years. The benchmark works only because the tasks behind it were written and checked by qualified people and the agents were run through one consistent model interface, the two layers his infrastructure exists to provide.
“A result that only works in your own lab is a marketing slide,” Wang notes. “We publish the benchmarks in the open because that is the only way to learn whether they hold up, and the only fair test of whether the infrastructure underneath them is any good.”
Where It Breaks
Building this kind of infrastructure is mostly a long fight with failure modes that stay hidden until scale. Recruiting experts in bulk is easy; holding the bar steady while doing it is not, because quality slips quietly long before the scores visibly drift. The model side has a mirror problem: the harder an interface is pushed, the more its weakest provider shows, and a single slow or inconsistent endpoint under heavy load can skew an entire evaluation without anyone noticing.
Wang's teams treat both as engineering problems rather than hiring ones. On the human side, every expert moves through the same pipeline of sourcing, vetting, training, and audited quality control, so the standard does not ride on who happened to pick up the work. On the model side, the access layer runs continuous checks on latency and consistency across providers and routes around any that degrade, so one unstable model cannot contaminate a batch of results. The lesson that keeps repeating, in his experience, is that quality matters far more than raw volume, and that the systems worth trusting are the ones designed to fail safely.
“Most teams underestimate how fast quality erodes when you scale people, and how much one unreliable model can wreck a benchmark,” Wang observes. “We put more effort into catching those failures early than into adding capacity, because a number you cannot trust is worse than having no number at all.”
The Cost of Building It Alone
The stakes land hardest on the institutions with the least room to spare. Frontier development has turned spectacularly expensive: training one leading model now reaches nine figures, with the compute behind a single flagship system, Gemini Ultra, estimated at roughly $191 million. Most universities and smaller labs will never train at that level, yet they are still expected to evaluate and build on the models that come out of it. The supporting infrastructure of expert review and multi-model access carries its own steep price when every group rebuilds it alone, and shared infrastructure is often what keeps those teams in the research at all.
This is the argument Wang keeps making for treating human expertise and model access as shared infrastructure rather than private overhead. When a smaller lab can reach the same vetted experts and the same standardized model interface as a frontier team, the thing that decides who gets to do serious work shifts back toward ideas and away from budgets. His plan is to keep widening that access, adding models and expert domains and pushing the onboarding and quality systems harder, on the bet that the next stretch of progress will come less from any single larger model and more from many more teams being able to test rigorously against the ones that already exist.
“The labs training the biggest models will keep making the headlines, and they have earned them,” Wang reflects. “But most of the useful research over the next few years will be done by people who could never afford to build any of this on their own. Giving them the same ground to stand on is the part of the work I care about most.”
1 Like
Comment

The push toward Industry 4.0 has made factory data central to modern manufacturing. The global smart manufacturing market reached $175 billion in 2025 and is projected to reach $274 billion by 2030, but many manufacturers still struggle with the same basic question: how do you turn machine data from many plants into information people can actually use?
Nabarun Bandyopadhyay is a Senior Data and AI professional and Sr. Delivery Consultant at Amazon Web Services, with 19 years of experience architecting enterprise scale data and AI platforms for large global organizations. In this interview, he explains why Industry 4.0 programs can stall after the vision is approved, what breaks inside the data foundation, and why repeatable plant onboarding is becoming one of the most important engineering problems in smart manufacturing.
Nabarun, thanks for joining us today. In simple terms, why do Industry 4.0 programs stall after the vision is approved?
The simple answer is that many manufacturers already have the ambition and the machine data. What they often do not have is the engineering foundation to bring that data together across different plants, machines, operating models, and reporting needs.
Think of it like trying to organize traffic from many busy roads into one shared system. Each plant is sending signals. Each machine has its own rhythm. If the foundation is not designed for that volume and concurrency, the data becomes hard to trust.
In one smart manufacturing data lake engagement, the issue was not lack of data. It was plant data moving at production speed from dozens of globally distributed plants, with multi-terabyte daily flows, millions of plant level records, and high frequency ingestion into shared analytical structures. At that point, a dashboard is not the hard part. The hard part is making sure the data arrives, is sequenced correctly, and can be governed without breaking the system.
That is also how I look at technical work when reviewing submissions for The 1st International Conference on Statistical Learning, Data Science and Generative AI (ICSL-DSGA 2026). A strong idea still has to survive real operating pressure. Industry 4.0 is no different.
What was the engineering gap nobody talks about?
The gap was concurrency. Multiple plants were writing high frequency data into shared transactional data lake tables, and the standard approach was not ready for the way those streams collided at production scale.
If you process everything one after another, the system becomes too slow. If you let everything write at the same time without control, you risk conflicts. So the problem was really about order, timing, and trust.
I designed a custom optimistic concurrency control mechanism with a queue based parallel processing model. The system assigned write positions just before the final data lake write, managed batch metadata, sequenced transformations, set job priorities, and prevented conflicting writes. In plain terms, it allowed plant data to move in parallel without letting those parallel processes corrupt the shared tables.
Why should factory leaders care about a concurrency problem?
Because they do not experience it as a concurrency problem. They experience it as slow failure visibility, incomplete machine status, delayed waste signals, or productivity reports that arrive too late to help.
The big data analytics market in manufacturing stood at $7.30 billion in 2025 and is forecast to reach $14.30 billion by 2030, which reflects how much manufacturers depend on factory data becoming useful intelligence. In this engagement, the concurrency design materially shortened the failure response window by moving from slower sequential patterns to govern parallel ingestion. It also supported visibility across multiple critical manufacturing KPI areas, including machine effectiveness, machine status history, productivity, and wastage.
Why is more custom development per plant the wrong answer?
Because custom development does not scale across a global plant network. Every plant has local differences, but the engineering model cannot be rebuilt from the beginning each time.
The better answer is metadata driven and configuration based. In this project, transformation logic was packaged into reusable components. That meant a plant could be onboarded by configuring metadata, sequencing rules, and transformation behavior instead of writing a separate solution for every site.
That changes the delivery model. The plant team still gets flexibility, but the engineering team is not trapped rebuilding the same pattern again and again. I also guided customer developers so the framework could be maintained and extended without every onboarding step returning to custom engineering.
How does this connect to sustainability and ESG outcomes?
Sustainability depends on measurement. If machine performance, waste, and productivity signals are late or inconsistent, ESG reporting becomes manual and reactive.
The smart manufacturing data lake helped contextualize machine performance data so leaders could track downtime, wastage, productivity, and machine status in a more consistent way. That same data foundation supported sustainability metric tracking by making factory performance more visible at plant level and across the broader operation.
What made the work useful beyond one customer engagement?
The useful part was repeatability. The concurrency solution solved an immediate blocker, but the metadata driven framework created a delivery pattern that could inform future enterprise data lake work.
That is important because many data programs repeat the same mistakes. They solve one problem once, then start again from zero on the next project. A better architecture captures the pattern so future teams can move faster with more structure. The strongest work is not just technically interesting. It can be maintained, repeated, and adapted under real constraints.
Where do data lakes fit into the future of Industry 4.0?
Data lakes are useful when they are designed for the workload. They are not useful when they become a place where data is stored but not shaped for real operations.
The global data lake market was valued at $11.07 billion in 2025 and is projected to reach $84.27 billion by 2034, but manufacturing use cases need more than storage growth. They need high throughput ingestion, transactional consistency, reusable transformation logic, and plant level onboarding that can support real factory needs.
In this project, the data lake supported machine effectiveness and machine status history use cases while helping the customer move from legacy ISA95 patterns toward Industry 4.0 standards. The goal was not simply to collect more data. It was to create a foundation that could support decisions closer to the machines.
What did stakeholders recognize in the project?
The response focused on execution under pressure. I was recognized for understanding the manufacturing business deeply, protecting delivery through a difficult technical workaround, and creating scalable solutions with long term impact.
That feedback points to the real lesson of the work. The challenge was not only technical. It required trust with customer teams, careful translation of plant level needs, and the ability to guide developers toward a maintainable framework. Strong architecture is not separate from delivery. It is how delivery survives complexity.
What should manufacturers take away from this?
The main lesson is simple: Industry 4.0 has to be engineered before it can be celebrated. Machine data, plant ambition, and executive sponsorship are not enough if the foundation cannot handle concurrency, throughput, onboarding, and sustainability measurement.
Manufacturers should ask practical questions early. Can the platform handle concurrent plant loads? Can new plants be onboarded through configuration? Can machine effectiveness, machine status, wastage, productivity, and ESG metrics be tracked from the same trusted foundation?
1 Like
Comment

For most of the programmatic era, the open internet has had structural disadvantages compared to the walled gardens.. Closed platforms know who their users are, what they watch, and what they buy, and they apply that knowledge at the exact instant an ad decision is made. The thousands of independent publishers that make up the rest of the web have rarely had anything comparable. The imbalance shows up in the money: walled gardens captured 78% of global digital advertising revenue in 2022, a share projected to reach 83% by 2027. The open question for the industry is whether that trajectory is fixed, or whether the open web can build the intelligence it has been missing.
Punit P. Shah, Director of Product Marketing at a global advertising technology platform, has spent over 15 years driving AI-powered products, platform strategy, and go-to-market execution across the technology and media ecosystem. Over the past year and a half, he led the strategy and worldwide rollout of an AI-powered programmatic curation platform built to give the supply side the real-time intelligence it has historically lacked. He was also a featured presenter at the 2025 Pollies Awards & Conference in Colorado Springs, where he spoke on winning with addressable media.
We spoke with Punit about why the open internet fell behind, what it takes to run decisioning inside a live auction, and why the next phase of programmatic advertising will be decided on the supply side.
The open internet has been losing ad share to closed platforms for over a decade. What do walled gardens actually have that independent publishers don't?
Signal and scale, applied at the moment of decision. That is the entire gap. A closed platform sees a logged-in user, their history, their context, and their likely response, and it uses all of that in the same instant it selects an ad. Nothing about that is magic. It is intelligence sitting next to inventory, with no distance between the data and the decision.
The open internet has always had the opposite architecture. Inventory lives with thousands of publishers, data lives with brands and data companies, and the decisioning lives inside buy-side platforms several hops away from the impression. Every hop loses signal and adds cost. So the open web ended up competing on volume and price while closed platforms competed on intelligence. That was never a fair fight, and the market share numbers reflect it.
You've spent the past year and a half rebuilding the supply side as an intelligence layer rather than a pipe. What does that look like in practice?
It starts with admitting that the traditional supply-side model was passive. We moved inventory, we ran auctions, and we left the thinking to everyone else. The platform we built to market reverses that. We unified inventory, data, identity, and algorithmic decisioning in a single system, so a curated deal is no longer a static list of approved sites. It is a live marketplace object that combines audience signals, segment qualification, contextual signals, and pricing logic, and it improves with every auction it touches.
The data partnerships were the foundation. We onboarded more than 300 data partners across audience, identity, and commerce signals, which lets buyers activate curated deals through any major demand platform without scattering their data across the ecosystem. My role covered the platform strategy, the product narrative, partner ecosystem development, and the global rollout across North America, Europe, and Asia Pacific. The hardest part was not the technology. It was getting an entire sales and partner organization to prove a supply-side platform as a place where decisions happen.
Walk me through the moment of the auction itself. Where does the decisioning actually live now?
Inside the auction, which is the whole point. We built a real-time infrastructure layer that lets partners deploy their own decisioning models directly into the auction path, so optimization happens while the bid request is live rather than hours later in a reporting loop. The supply side is the only position in the chain that sees the full picture at once: the impression, the page, the audience signal, and the price.
That position changes the math. When the decision happens next to the inventory, audience match rates run 30 to 40% higher than traditional demand-side targeting, and pricing efficiency improves by roughly 20% because buyers stop bidding blind. In campaign terms, public case studies on the platform documented outcomes like a 50% reduction in cost per acquisition. Those are not optimization rounding errors. They are the kind of gains that only appear when the intelligence moves to where the inventory is. This is only the starting point.
Real-time decisioning inside a live auction sounds unforgiving. What nearly broke?
Latency, first. An auction gives you milliseconds, and every model a partner wants to run has to fit inside that budget or the impression is gone. We had to be ruthless about what computes in real time and what gets precomputed upstream. The second problem was data quality. When you onboard several partners, you inherit hundreds of definitions of the same audience, and a curation platform that blends bad signals just industrializes waste.
And waste is exactly what we set out to eliminate. Only about 36 cents of every programmatic dollar reaches a publisher, with an estimated $26.8 billion lost to supply chain inefficiencies in 2025 alone. You do not fix that by adding another intermediary. You fix it by collapsing steps, which is what auction-level decisioning and audience qualification does. Fewer hops, fewer fees, less guessing. The discipline we forced on ourselves was that every layer of intelligence had to remove cost from the chain, not add a new toll booth to it.
You also spend time evaluating early-stage builders. What does that vantage point show you about where this field is heading?
Serving as a judge at the Builders of Tomorrow hackathon was a useful mirror. I assessed projects on whether the intelligence sat where the data actually lived, because that is the same architectural question my industry spent 15 years getting wrong. The strongest builders did not ask which platform to buy from. They asked where and how the decision should be made. That instinct is becoming the default for the next generation of engineers, and it should be.
Judging also reminded me how fast the tooling has moved. Teams stood up working decisioning prototypes in a weekend that would have taken a quarter to build a few years ago. The barrier to entry for intelligent systems has collapsed. What has not collapsed is the judgment about where to place and scale them, and that is what I look for, whether I am scoring a hackathon entry or reviewing a product roadmap.
Identifiers keep eroding and privacy regulation keeps tightening worldwide. Does that environment actually favor the supply side?
Yes, and I will say that plainly. The identifier era trained the industry to believe targeting is a buy-side job, because third-party cookies let demand platforms follow people across the web. That assumption is expiring. First-party data lives with publishers. Consent is collected by publishers. The durable signals that remain, context, attention, and commerce, sit closest to the supply side. The architecture is rebalancing toward where the data is born.
The platforms that survive the next 5 years will be the ones that embedd decisioning at the moment of the auction instead of bolting identity patches onto a model regulators are dismantling. We designed for that from the start: hashed identifiers, consent-based data, and first-party activation frameworks that reduce reliance on third-party cookies. Privacy law is not the threat to the open internet. Architecture built for a cookie world is.
What still has to happen for the open internet to genuinely close the gap?
Transparency and proof. Buyers need to see exactly what a curated deal costs, which signals power it, and what they received for the fee, because the open web cannot ask for budget on trust alone. The technology argument has largely been won. The accountability argument is still being made, and supply-side platforms are leading it rather than wait for buyers to force it.
The scale is coming either way. Programmatic is on track to account for 90% of worldwide display ad spending in 2026, which means the question is no longer just whether automation wins but whose intelligence runs it. My plan for the next 2 years is concrete: deepen the data partnerships, open the decisioning layer to more partner models, and keep publishing performance evidence in public case studies so the results can be inspected rather than asserted. As the open internet competes on intelligence instead of price, the gap with walled gardens stops being structural. It becomes a contest, and contests can be won.
1 Like
Comment
Most large enterprises cannot answer a basic question about their own infrastructure: how many APIs do they actually have. 78% of enterprise decision-makers admit they don't know how many APIs are running in their environment, and a separate survey found that 74% consider a significant share of those APIs unmanaged. The operational consequence is not theoretical. Fragmented visibility across gateway environments means incidents take longer to isolate, compliance gaps go undetected, and teams migrating between platforms have no reliable picture of the estate they are inheriting. The discipline of API governance has matured, but the infrastructure for putting it into practice at scale is still catching up.
Khushan Adatiya is a Senior Software Engineer with over a decade of experience building distributed systems at scale, spanning cloud infrastructure, API governance, and security. He led the design and delivery of API Insights within Google's API Hub platform, the first system to unify runtime analytics across all Apigee gateway deployment models in a single management interface. He is also the author of the published analysis "Scalability With AI: Lessons From Real Production Systems", which examines how production AI systems fail at scale and what engineering teams can do before the failure arrives.
We spoke with Adatiya about what it actually takes to give enterprises a unified view of their API landscape, and why that problem is harder than it sounds.
Large enterprises have been managing APIs for over a decade. Why is visibility still this hard to get right?
The visibility problem is not primarily technical. It is organizational. APIs get created by different teams, on different timelines, using whatever gateway or deployment model that team was already running. Some are on modern managed infrastructure. Some are on hybrid deployments. Some are on systems that have been running since before the cloud strategy was written. Over time, you accumulate an estate where no single person or team has the complete picture, because no single system was ever designed to provide it.
What makes it persistently hard is that proliferation happens faster than governance. A team ships a new integration, and by the time those endpoints are catalogued, new ones are already in flight. The management layer is always a sprint behind. You cannot govern what you cannot see, and in most large organizations the gap between what is deployed and what is tracked is larger than anyone wants to admit.
Your work on API Insights was specifically about building unified analytics across deployment models that were never designed to share data. What did that actually require?
It required treating each gateway type as a distinct data source with its own ingestion path, rather than assuming a common format and trying to normalize on the way in. Apigee X, Hybrid, Edge, and the on-premise deployment each produce runtime data differently. The ingestion architecture had to be adapted for each, and then the correlation logic had to reliably map that data back to the right API, version, and deployment entity inside the hub. That mapping sounds straightforward until you encounter cases where metadata is incomplete or where the same API exists in multiple environments at different versions. Getting that right under production conditions is the engineering problem, not the pipeline itself. The motivation is clear enough. 57% of organizations suffered an API-related data breach in the past two years, with the majority of incidents tracing back to endpoints that were either unmonitored or managed outside formal governance structures.
We also built the data pipeline to handle customers running multiple gateway types simultaneously, which is the common case in large enterprises. A company migrating from an older deployment to a newer one does not turn off the old system on day one. They run both in parallel, sometimes for years. The platform had to give them a coherent view of that mixed state, not just the endpoints that had already been modernized. That meant the analytics had to work across the full estate as it actually existed, not as a target architecture diagram imagined it should look.
What is the hardest engineering problem inside a multi-gateway analytics pipeline that nobody talks about publicly?
Data correlation at the entity level. When you have runtime traffic coming from four different gateway types, you need a reliable way to say: this call belongs to this API, this API belongs to this product, this product belongs to this team. That chain of attribution sounds obvious, but in practice it depends on identifiers being consistent across systems that were built by different teams with different conventions. When they are not consistent, and often they are not, you have to build a reconciliation layer that handles the gaps without losing data or introducing false attribution.
The other hard problem is security compliance within customer-controlled environments. Some enterprise customers operate inside VPC Service Controls, which restricts how data moves between services. Some require Customer Managed Encryption Keys. Each constraint changes how the pipeline must be architected, and you cannot simply route data through a shared processing layer and call it done. You have to design the pipeline to operate within those perimeters and then test it in network topologies that only surface in specific customer configurations. That work is invisible outside the team, but it determines whether the product is actually deployable in regulated industries.
When you were building this, what nearly caused the project to fail?
The security perimeter work almost broke the timeline twice. We had data flows that looked correct in our test environment and broke under VPC Service Controls in a customer environment, because the network topology differences changed which services could communicate with which. Diagnosing that required building out specific reproduction environments and working through ingress and egress rules across Cloud Run, Pub/Sub, Cloud Storage, and Cloud Monitoring simultaneously. It was not one misconfiguration. It was a chain of them, and each one you fixed revealed the next one downstream.
The second near-failure was grant delegation with Customer Managed Encryption Keys. When a customer uses CMEK, specific relationships between service accounts must be established for data to flow through an encrypted pipeline. We caught a gap in how we were handling delegation late enough that it required rearchitecting part of the ingestion layer under time pressure. That kind of problem is not in any documentation. You find it by running into it.
You have written publicly about where production AI systems break down. What does your experience building data infrastructure tell you about where AI at scale goes wrong most often?
The failure modes I have seen most consistently are not model failures. They are infrastructure failures that happen to involve a model. A system that worked well in testing degrades in production because the data distribution shifted and nobody had a monitor on the distribution. Or a pipeline that handled expected query volumes could not handle a 10x spike because the load testing assumptions were wrong. Or a service that was designed for synchronous request-response patterns gets asked to do something with a different latency profile and the timeouts are all wrong. None of these are problems specific to AI. They are the same problems that break any production system, and teams working on AI infrastructure often underestimate how much of the work is distributed systems engineering applied to a new workload.
The thing that I find most underestimated is the gap between what a model does in a benchmark and what it does when it is a dependency inside a larger system. A model that produces an answer 95% of the time is fine in isolation. Inside a system where five other components depend on that answer, a 5% failure rate compounds. Teams building production AI need to think about failure modes at the system level, not just at the model level. That means designing for graceful degradation, instrumenting the right signals, and testing the failure paths as carefully as they test the happy path.
You have judged engineering work across multiple technical venues. What separates submissions that hold up under scrutiny from the ones that do not?
The clearest signal is whether the team can explain what their system does when something goes wrong. Strong submissions have thought about failure modes specifically. They can tell you what happens when a dependency is slow, when data is missing, when load is three times what they designed for. Weaker submissions have usually only thought about the successful path. That gap shows up consistently across the work I have judged, including at DeveloperWeek 2026, where I served as a speaker and hackathon judge. The technical depth of a solution matters, but the quality of reasoning about its limits matters more.
The stakes for getting this right are not abstract. API attacks in the United States are projected to grow 548% by 2030. Teams building on API infrastructure now are building the systems that will need to absorb that pressure. The submissions that hold up under scrutiny are the ones that treated operational reliability and security as design constraints from the beginning, not as things to be added after the core functionality was working. That engineering discipline is what I look for, and it is what the industry is going to need more of.
If the next five years of investment in API governance goes exactly where it should, what does enterprise infrastructure look like in 2030? And what is the thing most likely to prevent that from happening?
In 2030, governance is not a separate discipline from deployment. The best-run organizations will treat an API as a managed artifact from the moment it is created: catalogued, monitored, attributed to an owner, and subject to lifecycle rules that automatically flag when it has been idle, undocumented, or operating outside its defined behavior. The analytics infrastructure to support that already exists in pieces. The missing piece is adoption, which is partly a tooling problem and partly a cultural one. Teams need to see governance as a condition of shipping, not as a compliance exercise that happens afterward.
The thing most likely to prevent that is velocity pressure. When shipping fast is the only metric that gets measured, cataloguing, documenting, and governing APIs gets deferred indefinitely. The deferral is invisible until it is not, usually when a breach happens or a migration fails because nobody can produce an accurate picture of what exists. The organizations in the best position in 2030 are the ones that decided now to treat visibility and governance as engineering work. That decision does not require new tools. It requires treating the maintenance of the API estate as seriously as the development of the APIs themselves.
1 Like
Comment
Pathology sits at the center of almost every cancer diagnosis, yet the infrastructure supporting it is under severe strain. The digital pathology market reached $1.31 billion in 2024 and is projected to hit $3.86 billion by 2033, a trajectory driven by a hard clinical reality: demand for pathology interpretation is growing faster than the workforce trained to provide it. For AI to help close any part of that gap, the models doing the work have to be trained on data that reflects what pathologists actually encounter.
Supriya Vijay is a Senior Software Engineer with over a decade of experience in machine learning and artificial intelligence. She is a co-author of the widely cited paper "Domain-Specific Optimization and Diverse Evaluation of Self-Supervised Models for Histopathology" and a named inventor on a pending U.S. patent application for self-supervised training of machine-learned image processing models for histopathology. She was a core technical contributor to research developing foundation models for analyzing microscopic tissue images, which have since been made publicly available as open research tools.
We spoke with Supriya about the challenge of adapting self-supervised learning to one of medicine's most demanding visual domains.
Pathology images are not like the images most AI systems are trained on. What makes them so different from an engineering perspective?
Histopathology images are fundamentally different from the kinds of images most computer vision systems are built to handle. A single whole-slide image can be 100,000 by 100,000 pixels or larger. The structures that pathologists examine - cell shapes, the spatial arrangement of nuclei, the density of the surrounding connective tissue - exist at multiple scales simultaneously. A model trained on images of everyday objects is not learning specifics that transfer meaningfully to this domain. The visual grammar is different.
The harder problem is that tissue appearance varies enormously across tissue types, staining protocols, and preparation methods - something well documented in the computational pathology literature. A slide prepared in one lab can have a completely different color profile than the same tissue prepared in another. Models trained on data from one institution often fail when deployed at a different one. That distribution shift is one of the central technical problems in the field, and it is why building a foundation model for pathology requires training strategies and augmentations engineered specifically for this domain.
Traditional AI development relies on massive amounts of hand-labeled data, which is a major bottleneck in medicine. How do foundation models change that equation?
The global density of pathologists is roughly 12.5 per million people, about 4 times below the North American level. Having those experts manually annotate massive training datasets is expensive, and the supply of people who can do it is limited. Self-supervised learning changes the equation - the model learns useful representations from unlabeled data by solving its own internal tasks, without requiring a pathologist to label every slide. The goal is to build a foundation model that learns rich, general-purpose representations across a wide range of tissue types and magnifications, which can then dramatically reduce the data, compute, and expertise needed to develop downstream tools.
Building a robust foundation model requires pre-training on a massive corpus of pathology images, optimizing the training strategy specifically for this domain, and evaluating the resulting representations across a highly diverse set of downstream tasks. That diversity of evaluation is essential - a model that performs well on one tissue type or one staining modality can still fail on others, so the evaluation suite must be structured to surface those failures. Today, high-quality pre-trained pathology models are increasingly available as open research tools. In the era of large language models and multimodal AI, these representations are especially critical - they serve as the foundational visual building blocks that allow language models to reason over massive gigapixel tissue slides alongside clinical text.
In self-supervised learning, the model has to learn what makes two images similar without human labels. How do you define what the model should pay attention to versus what it should ignore?
I think the important part is deciding what counts as a meaningful training signal when the ground truth is something only a trained pathologist can interpret. Self-supervised methods like SimCLR and Masked Siamese Networks learn by comparing different augmented views of the same image patch. The model is trained to recognize that a cropped, color-shifted, and blurred version of a tissue patch still represents the same underlying structure. But the augmentations that work well for natural images do not necessarily make sense for pathology. What counts as noise and what counts as signal are fundamentally different in this domain.
Getting that right requires careful thinking about how patches are sampled, how augmentations are designed, and which variations the model should learn to look past versus which it should preserve. That is where close collaboration with pathologists becomes essential. Staining variation, for instance, reflects differences in lab preparation, not in the underlying biology. Cellular density variation, on the other hand, carries information the model needs to retain. Those are competing pressures, and the pre-training design has to resolve them explicitly.
When engineers drill down from general computer vision into healthcare AI, what are the non-obvious traps they run into?
The most persistent challenge was getting the data partitioning right. In standard image benchmarks, you can split by image. In pathology, you have to split by patient or case, because there are often multiple slides from the same patient, and slides from the same patient can look nearly identical. If you split by slide instead of by patient, you risk data leakage. It is a well-known pitfall, but it introduces real complexity. You have to ensure that no patient's data appears in both the pre-training set and any downstream evaluation set, even across different tasks. Maintaining that separation across a benchmark spanning many tissue types and multiple data sources is nontrivial.
The other challenge is color. In standard AI training, you might randomly tweak an image's brightness or hue so the model learns to ignore changes in lighting. But in pathology, the color variation that matters does not come from lighting - it comes from differences in the chemical dyes used to stain tissue slides across different labs. Domain-specific stain augmentation methods, like RandStainNA, which simulates realistic variation in staining across different lab protocols, have proven to be among the most consistently beneficial improvements in this space. But the augmentation strategy has to be calibrated carefully: too aggressive and it destroys color information that pathologists say carries real diagnostic value in certain tissue types; too conservative and the model latches onto staining differences that are artifacts of preparation rather than biology. Finding that balance is iterative and requires clinical input throughout. That kind of domain expertise is not something you can substitute with more compute.
You have also served as a reviewer for both specialized health AI workshops and major machine learning conferences like ICML. What has reviewing taught you about the state of evaluation in the field?
Reviewing for the Structured Data for Health workshop is valuable because it enables you to read work from adjacent areas of health AI - clinical time series, EHR modeling, wearable data. The problems around data quality, distribution shift, and evaluation validity look structurally similar to what you would face in medical imaging. The specifics are different, but the failure modes are often the same: models that look good on benchmark splits but break on out-of-distribution cases, evaluations too narrow to surface the failures that would matter in deployment, training sets that over-represent well-resourced institutions.
That exposure reinforces the importance of rigorous evaluation design. It is easy to report a strong number on a single dataset and declare success. The harder and more informative thing is to test whether a learned representation holds up across tissue types, cancer subtypes, and preparation methods. That instinct came directly from seeing how often single-dataset evaluations fail to generalize.
Reviewing for ICML 2026 is a different kind of challenge. At the main conference, submissions span the full breadth of machine learning, including areas well outside my direct research. You have to assess rigor, novelty, and significance across work that might be in optimization theory, generative modeling, or reinforcement learning. That requires being honest about where your expertise ends and being careful not to over-weight what you happen to find interesting.
What reviewing at that scale has taught me is how much the quality of evaluation design varies even among strong submissions. The methods sections are usually tight. The evaluation sections are where the gaps show up.I look for those patterns in my own work now in a way I did not before.
The market for AI in pathology is expanding rapidly. As the field matures over the next few years, what do you see as the most critical unsolved engineering problems?
The hardest open problem is generalization across institutions. This is especially true as the field moves toward multimodal AI, pairing large language models with vision systems. An LLM can only reason effectively about an image if the underlying visual representation is robust. It is relatively tractable to get a model that works well on data from institutions that resemble your training set. It is much harder to build models that hold up when deployed in a setting with different scanners, different staining protocols, and different slide preparation practices. That gap between benchmark performance and real-world performance is where I believe the most important work still needs to happen.
The second challenge is workflow integration. You can build the most accurate model in the world, but if it requires pathologists to jump through extra software screens or does not integrate cleanly into existing workflows, it will not be adopted. Hospitals are dealing with fragmented IT infrastructure and massive data silos. Moving AI from a standalone research sandbox into an enterprise-grade platform is a massive, underappreciated engineering hurdle.
The final piece is evaluation infrastructure. It has become increasingly clear that we need evaluation datasets reflecting the actual diversity of pathology practice Building those datasets requires collaboration with clinical partners who have access to diverse, real-world tissue samples and the expertise to annotate them correctly. That is slow, careful work, and it does not attract the same attention as building a new model architecture. But it is foundational work. The AI in pathology market is projected to reach $347.4 million by 2030 at a 26.5% annual growth rate. If that investment funds new architectures but not the evaluation infrastructure to test them honestly, the field will keep producing work that looks stronger on benchmarks than it performs in practice. The research investment has to follow the hard problems, not just the publishable ones.
1 Like
Comment

Comment