Customers find out first
The site is down for forty minutes until somebody writes to support.
I add checks and alerts — your systems do the telling.
I am Dmitry Savichev. I make sure failures reach you from your systems rather than from your customers, and that a release does not depend on anyone remembering the right order of commands.
Remote, working solo. I usually reply the same day.
Dmitry Savichev
SRE/DevOps engineer
Sound familiar?
The site is down for forty minutes until somebody writes to support.
I add checks and alerts — your systems do the telling.
Twenty messages an hour, and the one that mattered is lost among them.
I separate the signals: one failure, one message.
A contractor set it up two years ago. Now everyone is afraid to touch it.
A written review of what you have and a plan for what to fix first.
Services
The work
All of it on my own infrastructure, which I run as production. There are no client projects here, and I do not pass mine off as anyone else's.
Server state and incidents open with a button in the chat.
Running in production since 31 August
Write-up
Journals had grown to 4 GB and kept going.
4.0 GB → 31 MB, no restarts
Write-up
The observation host lost connectivity and declared five healthy sites down.
A write-up of my own mistake
Write-up
Work ends with a report: what I found, what I did, how I verified it and what is left. Here is an example.
A couple of sentences about your infrastructure and what worries you is enough for me to tell whether I can help.