online Bengaluru, India IST (UTC+5:30)
Sameer Alam
Infrastructure & Security Engineer
I build, automate and operate infrastructure that is observable, reproducible, secure, diagnosable and resilient.
- Kubernetes
- Linux
- AWS
- Terraform
- Ansible
- CI/CD & GitOps
- Observability
- SRE
- DevSecOps
- Python · Go · Shell
- AI/LLM tooling
Quick commands
- building & operating infrastructure
- 8 yrs
- servers supported across RHEL, AIX, SUSE, Ubuntu
- 500 +
- infrastructure & automation projects built
- 6
- collection on Ansible Galaxy
- 1
$ ./healthcheck --principles
Automation-first, reliability-focused
Eight years building, automating and operating software and infrastructure — from enterprise Linux fleets to Kubernetes platforms.
I work automation-first: if a task happens twice, it becomes a script, a playbook or a pipeline. Then I make it observable so the next failure explains itself.
- observable
Every system emits the metrics, logs and events needed to answer “what is it doing right now?” without SSH.
- reproducible
Infrastructure is code. Any environment can be rebuilt from a repo and a pipeline, not from memory.
- secure
Least privilege, scanned dependencies and no secrets in source — security is a default, not a review step.
- diagnosable
Failures leave a trail. Runbooks, health checks and clear signals shorten the path from alert to root cause.
- resilient
Systems expect failure: they degrade gracefully, recover automatically and are tested against breaking.
$ kubectl get skills --all-namespaces
Stack
What I use to build, ship, observe and secure systems.
- ns/orchestration 6
Containers & orchestration
- Kubernetes
- OpenShift
- Docker
- Podman
- kind
- kubeadm
- ns/systems 7
Linux & systems
- RHEL
- SUSE
- Ubuntu
- Debian
- AIX
- Bash
- Networking
- ns/cloud-iac 4
Cloud & infrastructure as code
- AWS
- Terraform
- Ansible
- Ansible Galaxy collections
- ns/delivery 4
CI/CD & GitOps
- GitHub Actions
- GitOps
- Jenkins
- Git
- ns/observability 6
Observability & SRE
- Prometheus
- Grafana
- Nagios
- Root cause analysis
- Runbooks
- Incident response
- ns/security 5
Security
- DevSecOps
- Linux hardening
- TLS certificate lifecycle automation
- Secret scanning
- Least-privilege CI
- ns/languages 4
Languages
- Python
- Go
- Shell
- TypeScript
- ns/ai 6
AI / LLM tooling
- Ollama
- Hugging Face
- FastAPI
- Gradio
- RAG
- Gemini Enterprise
$ ls -l ~/projects
Projects
Tools I built to remove toil: remediation, fleet health, network visibility and AI-assisted automation.
- KubeRescue active
KubeRescue
Kubernetes failure detection and auto-remediation
Autonomous Kubernetes failure detection and policy-driven auto-remediation engine for SRE teams.
- Go
- Kubernetes
- SRE
- Incident response
★ 3 · last commit 3 weeks ago
read the deep dive on KubeRescue - linux-vitals on Ansible Galaxy
LinuxVitals
Agentless Linux fleet health checks
Ansible collection for Linux fleet health checks across RHEL, Fedora, Ubuntu and SUSE: baseline/postcheck comparison, opt-in self-healing and a self-contained HTML dashboard.
- Ansible
- Python
- Linux
- Monitoring
★ 3 · last commit 4 days ago
read the deep dive on LinuxVitals - ipmg v3.0 in progress
IPMG
Network monitoring from the command line
Python tool that finds which hosts on your network are up and what changed since last time: parallel ping sweeps, reverse DNS, scan history and diffs, Excel/CSV/JSON/Markdown reports and a local web UI.
- Python
- Networking
- CLI
★ 10 · last commit 3 days ago
read the deep dive on IPMG - private repo active
AI Ansible Generator
Plain English → validated Ansible playbooks
Turns a plain-English request into an Ansible playbook, then validates YAML, syntax and safety and runs a repair loop until it passes. FastAPI + Gradio, backed by Ollama or Hugging Face models.
- Python
- FastAPI
- Gradio
- Ollama
- Hugging Face
- Ansible
- private repo v2 rebuild
kubernetes-platform
Production-oriented Kubernetes platform
A production-oriented Kubernetes platform, currently being rebuilt from the ground up as v2.
- Kubernetes
- GitOps
- Terraform
- Observability
- private repo active
trend-publisher
Automated content pipeline
Automated content pipeline that publishes through the Meta Graph API.
- Python
- Meta Graph API
- Automation
$ ls ~/projects/more
- k8s-kubeadm-lab Reproducible multi-node kubeadm lab
- devops-case-studies Production case studies
- devterm One-command terminal setup for SREs
- SecureExamPortal Online exam platform with integrity controls
- llm-dev-kit Local LLM + RAG toolkit
$ gh api users/sameeralam3127
Live from GitHub
Pulled from the GitHub API on every deploy and nightly. Refresh to read the current numbers straight from GitHub.
- public repositories
- 14
- stars across repositories
- 54
- forks
- 10
- contributions in the last 12 months
- 3,270
View as table
| Week of | contributions |
|---|---|
| 27 Sept 2026 | 117 |
| 20 Sept 2026 | 125 |
| 13 Sept 2026 | 94 |
| 6 Sept 2026 | 106 |
| 30 Aug 2026 | 65 |
| 23 Aug 2026 | 53 |
| 16 Aug 2026 | 61 |
| 9 Aug 2026 | 50 |
| 2 Aug 2026 | 42 |
| 26 Jul 2026 | 138 |
| 19 Jul 2026 | 41 |
| 12 Jul 2026 | 59 |
| 5 Jul 2026 | 41 |
| 28 Jun 2026 | 49 |
| 21 Jun 2026 | 41 |
| 14 Jun 2026 | 33 |
| 7 Jun 2026 | 26 |
| 31 May 2026 | 36 |
| 24 May 2026 | 33 |
| 17 May 2026 | 28 |
| 10 May 2026 | 35 |
| 3 May 2026 | 51 |
| 26 Apr 2026 | 40 |
| 19 Apr 2026 | 62 |
| 12 Apr 2026 | 60 |
| 5 Apr 2026 | 70 |
| 29 Mar 2026 | 55 |
| 22 Mar 2026 | 50 |
| 15 Mar 2026 | 27 |
| 8 Mar 2026 | 33 |
| 1 Mar 2026 | 35 |
| 22 Feb 2026 | 62 |
| 15 Feb 2026 | 50 |
| 8 Feb 2026 | 63 |
| 1 Feb 2026 | 40 |
| 25 Jan 2026 | 52 |
| 18 Jan 2026 | 52 |
| 11 Jan 2026 | 71 |
| 4 Jan 2026 | 59 |
| 28 Dec 2025 | 61 |
| 21 Dec 2025 | 61 |
| 14 Dec 2025 | 94 |
| 7 Dec 2025 | 47 |
| 30 Nov 2025 | 88 |
| 23 Nov 2025 | 69 |
| 16 Nov 2025 | 51 |
| 9 Nov 2025 | 154 |
| 2 Nov 2025 | 84 |
| 26 Oct 2025 | 92 |
| 19 Oct 2025 | 45 |
| 12 Oct 2025 | 103 |
| 5 Oct 2025 | 64 |
| 3 Oct 2025 | 52 |
- sameeralam3127.github.io TypeScript
- stars
- 3
- forks
- 0
- open issues
- 0
last commit · Merge pull request #8 from sameeralam3127/feat/scroll-tint
- compute-central-docs HTML
Practical DevOps, cloud, Kubernetes, automation, AI engineering, and SRE documentation for Compute Central.
- stars
- 5
- forks
- 0
- open issues
- 0
last commit · Answer the searches that land on pages without the answer (#27)
- ipmg Python
Find out which hosts on your network are up, and what changed since last time. Parallel ping sweeps, reverse DNS, scan history and diffs, Excel/CSV/JSON/Markdown reports, and a local web UI.
- stars
- 10
- forks
- 4
- open issues
- 32
last commit · chore(release): 3.1.2
- linux-vitals Python
LinuxVitals — agentless Ansible collection for Linux fleet health checks (RHEL/Fedora/Ubuntu/SUSE): baseline/postcheck comparison, opt-in self-healing, and a self-contained HTML dashboard.
- stars
- 3
- forks
- 6
- open issues
- 30
last commit · deps: bump the dev-tooling group with 2 updates (#100)
- devterm Shell
One-command iTerm2 + macOS Terminal.app + Starship setup for developers and SREs — themed profiles, Nerd Font icons, and a git/Kubernetes/AWS-aware prompt.
- stars
- 1
- forks
- 0
- open issues
- 4
last commit · Show Mac name, CPU and RAM in the prompt on Terminal.app (#6)
- homebrew-tap Ruby
- stars
- 1
- forks
- 0
- open issues
- 0
last commit · Merge pull request #4 from sameeralam3127/dependabot/github_actions/g…
- devops-case-studies JavaScript
Practical DevOps, SRE, Cloud and Platform Engineering case studies covering architecture, troubleshooting, automation, reliability, security and production operations.
- stars
- 3
- forks
- 0
- open issues
- 12
last commit · Make the home page a landing page, and fix what the rename broke (#114)
- ansari Python
ANSARI
- stars
- 2
- forks
- 0
- open issues
- 6
last commit · docs: shorten the README and cut the docs to two (#35)
- SecureExamPortal Python
Production-oriented online exam platform: FastAPI + PostgreSQL API, React/Vite frontend, background worker, and Nginx edge. MCQ authoring, role-based dashboards, secure auth (password + Google), and exam-integrity controls.
- stars
- 7
- forks
- 0
- open issues
- 0
last commit · Bump nanoid from 3.3.16 to 3.3.18 in /frontend (#71)
$ pagerduty incident ack --on-call=you
Incident simulator
You're on call. Pick a page, read the evidence, and decide what to do next. Every step costs time on the clock, and each choice is reviewed the way I'd review it in a postmortem.
Pods in CrashLoopBackOff after a release
ALERT KubePodCrashLooping: payments-api pods restarting repeatedly in namespace payments. Checkout error rate above SLO.
HTTPS failing at the edge
ALERT Synthetic check FAILED for https://api.example.com/health: SSL certificate problem. Partner API error rate 41%.
$ cat .github/workflows/*.yml | less
How this site ships
The site is its own case study: every change goes through the same gates I'd put in front of production. Select a step to see what it does.
-
01 · trigger
-
02 · verify
-
03 · build
-
04 · audit
-
05 · deploy
-
06 · observe
trigger · main · pull_request
git push
Pull requests run the full CI suite; merges to main deploy. Nothing reaches production without passing the same checks.
All steps as text
-
1. trigger
- git push: Pull requests run the full CI suite; merges to main deploy. Nothing reaches production without passing the same checks.
- nightly cron: A nightly scheduled run rebuilds the site so GitHub stats and the latest Compute Central articles stay fresh without a commit. It can also be run by hand with workflow_dispatch.
-
2. verify
- lint · format · types: ESLint, a Prettier check and strictest TypeScript (including exactOptionalPropertyTypes and noUncheckedIndexedAccess) run on every pull request.
- tests: Vitest covers the data and logic: terminal commands, the incident engine, content invariants. Playwright smoke tests load every page, drive the terminal and fail on any console error.
- security: Secret scanning, a dependency audit that fails on high severity, and CodeQL for JavaScript/TypeScript. All actions are pinned to full commit SHAs and kept current by Dependabot.
-
3. build
- fetch data: Build-time scripts pull repo stats and the contribution calendar with the workflow's GITHUB_TOKEN, and the newest Compute Central articles from its sitemap. Results ship as static JSON, so visitors never hit API rate limits. Any failure falls back to curated content instead of failing the build.
- astro build: Astro renders every page to static HTML. Only the interactive pieces (terminal, incident simulator, command palette) hydrate as React islands; everything else ships zero JavaScript.
-
4. audit
- lighthouse ci: Lighthouse CI audits the built site and enforces a hard budget of 95 for Performance, Accessibility, Best Practices and SEO. The median scores are written into build-meta.json and shown in the status panel.
- link check: lychee checks every internal and external link in the built HTML so a renamed repo or moved article can't ship as a 404.
-
5. deploy
- github pages: The dist/ folder is uploaded as a Pages artifact and deployed to the github-pages environment with OIDC. Permissions are least-privilege (contents: read, pages: write, id-token: write) and a concurrency group cancels stale deploys.
-
6. observe
- post-deploy check: After deploy, the pipeline requests the live URL and fails the run if it doesn't return 200. The deploy success rate in the status panel comes from these runs.
$ kubectl get --raw /healthz?verbose
System status
This site’s own telemetry. The CI pipeline writes these numbers at build time, so they’re as current as the last deploy.
main · workflow run
3 Oct 2026, 14:08 UTC
—
data fetch → astro build
100%
4/4 recent runs
- Performance —
- Accessibility —
- Best practices —
- SEO —
Scores are recorded by Lighthouse CI during deploy.
$ git log --oneline career
Experience
8 years across enterprise support, systems engineering and infrastructure automation.
-
Jun 2021 → present
Technical Support Professional · IBM
- Operate and support 500+ servers across RHEL, AIX, SUSE, Ubuntu, Debian and Windows Server.
- Automate patching, health checks and routine operations with Shell, Python and Ansible.
- Built and maintain automation for SSL/TLS certificate lifecycle management.
- Root cause analysis using Prometheus, Grafana, Nagios, API monitoring and system-level debugging.
- Contribute CI/CD automation with GitHub Actions, plus runbooks and operational guidance for the team.
-
Feb 2020 → May 2021
System Engineer · EY
- Supported enterprise PHP and Java applications; diagnosed production issues and shipped fixes that cut downtime by 30%.
- Managed production releases, configuration changes and hotfixes before CI/CD was in place.
-
Dec 2018 → Feb 2020
System Engineer · Netsoft Consulting Services
- Maintained enterprise applications and operating systems; troubleshot production issues to keep downtime minimal.
-
May 2018 → Dec 2018
Solutions Engineer · Quatrro
- Remote technical support for US customers across network, OS and application issues.
-
2024
MCA
Jain University, Bangalore
-
2017
BCA
Magadh University
$ cat /etc/credentials
Open source & certifications
- IBM/docling-pipelines merged
Refactored the Ollama client to hoist repeated imports out of hot-path methods.
- K8sGPT (CNCF) contributing
Contributing to the CNCF project that diagnoses Kubernetes clusters with AI.
-
Google Cloud Certified Partner Specialist — Gemini Enterprise Agent Development
Google Cloud
-
Certified Partner Specialist — Gemini Enterprise Deployment
Google Cloud
-
Build with Gemini
Google
11 earlier credentials
- Monitor and Manage Google Cloud Resources (skill badge) · Google Oct 2024
- The Basics of Google Cloud Compute (skill badge) · Google Oct 2024
- AWS Knowledge: Cloud Essentials · AWS Jan 2024
- AWS Knowledge: Architecting · AWS Oct 2023
- Pen Testing, Incident Response & Forensics · IBM Sep 2023
- Networking Essentials · Cisco Feb 2023
- AWS Cloud Quest: Cloud Practitioner · AWS Dec 2022
- Docker Essentials: A Developer Introduction · IBM Nov 2022
- Cybersecurity Essentials · Cisco Nov 2020
- CyberArk Certified Level 1: Trustee · CyberArk Oct 2020
- NSE 2 Network Security Associate · Fortinet May 2020 (expired)
$ curl -s computecentral.in
Latest from Compute Central
Practical guides on Kubernetes, Ansible, Terraform, CI/CD, SRE and security, written from day-to-day operations work.
- #ansible updated 2 Oct 2026 Ansible Inventory for Dev, Staging, and Prod Environments How to structure Ansible inventory for dev, staging, and production — separate inventories, per-environment group_vars, and guardrails so a mistake can
- #cloud updated 2 Oct 2026 AWS CLI for Accounts and Organizations: Account ID, SSO, SCPs Get your AWS account ID with aws sts get-caller-identity, run AWS Organizations and IAM Identity Center from the CLI, plus SSO profiles, SCPs, and budgets.
- #kubernetes updated 2 Oct 2026 Kubernetes Debugging: Events, describe, and exec A repeatable Kubernetes debugging methodology using kubectl describe, get events, ephemeral debug containers, exec, port-forward, and cp.
- #foundations updated 26 Sept 2026 How Git Works: Commits, Trees, Blobs, Refs, and HEAD Build a mental model of Git internals — the object database, commits as snapshots, the staging area, branches as movable pointers, HEAD, and remotes.
- #ai-guide updated 24 Sept 2026 LLM Fundamentals: Tokens, Context Windows, and Prediction How an LLM works: it predicts the next token from everything before it, one token at a time. Tokens, context windows, probability, and why models hallucinate.
- #cicd updated 24 Sept 2026 Code Quality: SonarQube, Linters, and Quality Gates Build automated code quality checks — open-source linters and scanners, SonarQube installation and quality gates, and CI pipelines that enforce them.
$ ansible-playbook hire-me.yml
Contact
Questions about a project, a guide I wrote, or infrastructure and security work. Email is the fastest route.