Written by people who run this in production
31 articles across six clusters. Technical, specific, and willing to say when a popular practice is not worth the trouble.
What does the DevOpsArk blog cover?
The DevOpsArk blog publishes 31 articles on DevOps, Kubernetes, observability, AI DevOps, DevSecOps and cloud, written by the engineers who build the platform.
Start with a subject
DevOps
Fundamentals, automation strategy, incident practice and the platform-versus-toolchain question.
Kubernetes
Architecture, monitoring, deployment strategy, multi-cluster operations and troubleshooting.
Observability
Metrics, logs, traces, alerting design and what observability actually means.
AI DevOps
Agentic DevOps, AI log analysis, incident response and where automation should stop.
DevSecOps
Container and Kubernetes security, vulnerability management, secrets handling, audit trails and compliance evidence.
Cloud and cost
Multi-cloud operations, cost optimisation, infrastructure drift and infrastructure hygiene.
Everything, newest first
Blameless postmortems: the format that actually gets filled in
Most postmortem templates go unused after the first few sections. Why that happens, and what a blameless postmortem format looks like when people actually complete it.
Compliance readiness vs compliance certification: what auditors actually ask for
"Compliance ready" and "certified" get used interchangeably, but auditors treat them very differently. What the difference is, and what audit evidence actually looks like.
What is infrastructure drift, and why your Terraform state keeps lying to you
Infrastructure drift is the gap between what your Terraform state says and what is actually running. Why it happens, what breaks, and how to catch it before an incident.
SIEM vs audit log vs observability: three different jobs, constantly confused
SIEM, audit logs and observability all involve data about your systems, but they answer different questions for different people. How to tell them apart.
DevOps audit trails: why full event history is your cheapest incident tool
Most audit trails are built for a compliance reviewer and used, months later, by an on-call engineer at 2am. What a useful trail captures and how to make it serve both jobs.
SSL certificate management: preventing the outage nobody planned for
Why certificate expiry still causes outages, how to find the certificates nobody documented, and how to automate renewal so the problem stops recurring.
Kubernetes troubleshooting: a decision tree that works
The common Kubernetes failures, what each one actually means, and the order of checks that reaches the cause fastest.
Multi-cloud DevOps: making one operating model work across providers
Why organisations end up multi-cloud, what it genuinely costs, and how to build one operating model across providers without pretending they are identical.
AI versus traditional automation: choosing the right one
A decision framework for when a deterministic script is the correct answer and when a reasoning agent earns its extra complexity and risk.
Secrets management: getting credentials out of your repositories
How to centralise credentials, replace long-lived secrets with short-lived ones, rotate without breaking consumers, and respond when a secret leaks.
AI-powered incident response: compressing the first ten minutes
Where incident time actually goes, which parts an agent can take over safely, and how to structure incident response so the automation helps rather than adds noise.
AI log analysis: making millions of lines legible
How log pattern extraction works, why it finds errors that search never will, and how to use novelty and rate detection without generating a new source of noise.
Vulnerability management: turning a report into a work queue
How to make a twelve-thousand-row vulnerability report actionable: deduplication, exposure-based ranking, ownership and verified closure.
DevOps platform or toolchain? An honest comparison
The real trade-off between assembling best-of-breed tools and adopting an integrated platform, including the costs of each that vendors on both sides tend not to mention.
Alerting that people still read at 3am
How to build an alerting system with a high proportion of actionable pages: correlation, ownership routing, suppression, and deleting the rules that never produce a decision.
AI agents in DevOps: where they help and where they do not
A practical assessment of where AI agents genuinely improve infrastructure work, where they are oversold, and how to introduce them without creating a new class of incident.
Kubernetes deployment strategies: rolling, blue/green and canary
How rolling updates, blue/green and canary deployments differ, what each actually protects you from, and how to choose per environment rather than per opinion.
Kubernetes security: the controls that matter most
A prioritised guide to securing Kubernetes (RBAC, pod security, network policy, secrets and supply chain) ordered by risk reduced rather than by chapter number.
Logs, metrics and traces: which signal answers which question
What each telemetry type is genuinely good at, what it costs, and how to decide where a given piece of information belongs.
Kubernetes cost optimization: where the money actually goes
Why Kubernetes clusters cost more than they should, how to attribute spend to teams, and the specific changes that produce the largest savings.
Container security: the practices that actually reduce risk
Build-time hardening, runtime restriction and supply chain controls for containers, ordered by how much risk each one removes rather than by how often it is mentioned.
How to automate Docker builds without hand-writing Dockerfiles
Multi-stage builds, layer caching that actually works, hardening defaults, and how to generate and maintain container definitions across a large service estate.
Monitoring vs observability: a distinction worth keeping
The difference between monitoring and observability, why the distinction is more than marketing, and what each one is actually for.
Multi-cluster Kubernetes management: patterns that hold up
Why organisations end up with many Kubernetes clusters, what actually gets hard at that point, and the operational patterns that survive contact with a growing fleet.
DevOps automation: what to automate, in what order
A practical sequence for automating a delivery path, which step gives the most back first, which automation tends to be regretted, and how to tell when a stage is genuinely done.
What is DevSecOps? Beyond "shift left"
What DevSecOps means in practice, why shifting left fails when the feedback is not actionable, and the practices that actually change security outcomes.
Kubernetes monitoring: what to measure and why
A practical guide to Kubernetes monitoring: the cluster, node and workload signals that predict failure, the ones that only look useful, and how to alert on them without drowning.
What is observability, and how is it different from monitoring?
A definition of observability that does more work than "the three pillars": what property you are actually trying to obtain, and how to tell whether you have it.
What is agentic DevOps?
A precise definition of agentic DevOps, how it differs from scripted automation and from AIOps, and the conditions under which an agent is safe to give real access.
Kubernetes architecture explained: the parts that matter operationally
What each Kubernetes component actually does, which ones you will meet during an incident, and the mental model that makes cluster behaviour predictable.
What is DevOps? A definition that survives contact with practice
What DevOps actually means, where the definition came from, what the lifecycle looks like in practice, and the common misreadings that turn it into a job title instead of a way of working.
Who writes here
Harshit Sengar
Harshit works on the DevOpsArk control plane and writes about Kubernetes operations, agentic automation and the practical economics of running infrastructure at scale. Most of what appears here comes out of production incidents rather than reading.
DevOpsArk Engineering
Articles written collectively by the DevOpsArk engineering team, the people who build the modules described on this site. Technical deep dives, architecture notes and the reasoning behind specific product decisions.
DevOpsArk Security
The DevOpsArk security engineering group writes about DevSecOps practice, container and Kubernetes security, vulnerability prioritisation and the parts of compliance that actually change how systems are built.
Questions about the blog
The DevOpsArk blog publishes 31 articles on DevOps, Kubernetes, observability, AI DevOps, DevSecOps and cloud, written by the engineers who build the platform.
The DevOpsArk engineering and security teams, plus named individual authors. Every article has an author page listing their expertise and other articles.
Each article covers its topic on its own terms and mentions DevOpsArk where it is genuinely relevant, usually in one clearly marked section. If an article is only useful to someone already using the product, it has failed.
Articles carry both a published and an updated date. Technical material is revised when the underlying behaviour changes rather than on a schedule.
Yes, at /rss.xml. It includes every article with its cluster, author and publication date.
Prefer a structured route through this?
The learning tracks arrange these articles into ordered paths per discipline.