Netomi’s lessons for scaling agentic systems into the enterprise
How Netomi scales enterprise AI agents using GPT-4.1 and GPT-5.2—combining concurrency, governance, and multi-step reasoning for reliable production workflows.
Insights
Practical writing on software architecture, SaaS products, AI automation, legacy modernisation, and the business of building reliable systems.
Curated links from external sources — not 360Softy original articles.
How Netomi scales enterprise AI agents using GPT-4.1 and GPT-5.2—combining concurrency, governance, and multi-step reasoning for reliable production workflows.
Cloudflare admin activity logs now capture each time a DNS over HTTP (DoH) user is created. These logs can be viewed from the Cloudflare One dashboard ↗, pulled via the Cloudflare API, and exported through Logpush.
If you maintain a CSI driver that uses service account tokens, Kubernetes v1.35 brings a refinement you'll want to know about. Since the introduction of the TokenRequests feature, service account tokens requested by CSI drivers have been passed to them through the volume_context field. While this has worked, it's not the ideal place for sensitive information, and we've seen instances where tokens were accidentally logged in CSI drivers. Kubernetes v1.35 introduces a beta solution to address this
2025 was a defining year for DigitalOcean, not only because we shipped more products and features than ever before, but because we solidified our vision about what the next era of cloud and AI will look like. We supported customers as they ran inference at scale, launched new products, engaged with our community in-person and online, and built out our inference cloud, which gives digital native enterprises and AI-native businesses the power to integrate AI and cloud workflows through one unified
Last year we introduced the , and described how the v0 models operate inside a multi-step agentic pipeline. Three parts of that pipeline have had the greatest impact on reliability. These are the dynamic system prompt, a streaming manipulation layer that we call “LLM Suspense”, and a set of deterministic and model-driven autofixers that run after (or while!) the model finishes streaming its response.v0 Composite Model Family What we optimize for The primary metric we optimize for is the percenta
We open-sourced , the Bash execution engine used by to reduce our token usage, improve the accuracy of the agent's responses, and improve the agent's overall performance.bash-toolour text-to-SQL agent that we recently re-architected gives your agent a way to find the right context by running bash-like commands over files, then returning only the results of those tool calls to the model.bash-tool Context windows can fill up quickly if you include large amounts of text into a prompt. As agents t
Teams can now create, update and delete Secure Compute networks directly from the Vercel dashboard, the API, and Terraform. Secure Compute networks provide private connectivity between your Vercel Functions and backend infrastructure and let you control regional placement, addressing, egress and failover of your projects. Now you can: This is available today for Enterprise teams. to get started.Check out the documentation Read more with no contract amendment or manual provisioning
Tolan built a voice-first AI companion with GPT-5.1, combining low-latency responses, real-time context reconstruction, and memory-driven personalities for natural conversations.
Work with 360Softy
Book a free consultation and we will tell you honestly whether we can help.