360SOFTY

Insights

Engineering Insights

Practical writing on software architecture, SaaS products, AI automation, legacy modernisation, and the business of building reliable systems.

RSS

Curated links from external sources — not 360Softy original articles.

ExternalCloud
Cloudflare Changelog

Cache - Asynchronous stale-while-revalidate

Cloudflare's stale-while-revalidate support is now fully asynchronous. Previously, the first request for a stale (expired) asset in cache had to wait for an origin response, after which that visitor received a REVALIDATED or EXPIRED status. Now, the first request after the asset expires triggers revalidation in the background and immediately receives stale content with an UPDATING status. All following requests also receive stale content with an UPDATING status until the origin responds, after w

Cache
Cloudflare ChangelogRead original
ExternalCloud
Cloudflare Changelog

Cloudflare Fundamentals - Markdown responses for Cloudflare 1xxx errors

Cloudflare now returns structured Markdown responses for Cloudflare-generated 1xxx errors when clients send Accept: text/markdown. Each response includes YAML frontmatter plus guidance sections (What happened / What you should do) so agents can make deterministic retry and escalation decisions without parsing HTML. In measured 1,015 comparisons, Markdown reduced payload size and token footprint by over 98% versus HTML. Included frontmatter fields: error_code, error_name, error_category, http_sta

Cloudflare Fundamentals
Cloudflare ChangelogRead original
ExternalBackend Development
FastAPI Releases

0.133.1

Features 🔧 Add FastAPI Agent Skill. PR #14982 by @tiangolo. Read more about it in Library Agent Skills. Internal ✅ Fix all tests are skipped on Windows. PR #14994 by @YuriiMotov.

FastAPI ReleasesRead original
ExternalAI
NVIDIA Technical Blog

Making Softmax More Efficient with NVIDIA Blackwell Ultra

LLM context lengths are exploding, and architectures are moving toward complex attention schemes like Multi-Head Latent Attention (MLA) and Grouped Query... LLM context lengths are exploding, and architectures are moving toward complex attention schemes like Multi-Head Latent Attention (MLA) and Grouped Query Attention (GQA). As a result, AI ”speed of thought” is increasingly governed not by the massive throughput of matrix multiplications, but by the transcendental math of the softmax function.

NVIDIA Technical BlogRead original
ExternalCybersecurity
Google Security Blog

Staying One Step Ahead: Strengthening Android’s Lead in Scam Protection

Posted by Lyubov Farafonova, Product Manager, Phone by Google; Alberto Pastor Nieto, Sr. Product Manager Google Messages and RCS Spam and Abuse shared how Android’s proactive, multi-layered scam defenses utilize Google AI to protect users around the world from over 10 billion suspected malicious calls and messages every month1. While that scale is significant, the true impact of these protections is best understood through the stories of the individuals they help keep safe every day. This incl

Google Security BlogRead original

Work with 360Softy

Building a SaaS product, AI system, or business platform?

Book a free consultation and we will tell you honestly whether we can help.