Site Reliability Operations with ServiceNow - SRE Mastery
hello everyone my name is michel i'm the lead strategic advisor at aina and partners and today we are looking at one of servicenow's newest solution in the shared itsm incident and itom event management space i am talking about site reliability operations srops service now is solution for businesses driven by or switching to site reliability engineering sre [Music] for devops driven practices sre adoption will only continue to grow the practice and culture shift will take priority in 2021 in other words more people not only sres will have a reliability mindset approaches such as observability run book automation and blameless retrospectives will become deeply rooted in business practices devops is one could argue similar to itil more like a concept of framework of practices however it is often viewed as more theoretical sre on the other hand i quote wikipedia in 2003 google developed site reliability engineering sre an approach for releasing new features continuously into large scale high availability systems while maintaining high quality and user experience while sre predates the development of devops they are generally viewed as being related to each other although from unreliable sources according to wikipedia policy one can safely assume sre to be an in real life application of the devops framework even if sre predates devops then for my next reasoning you look at cre customer reliability engineering well what is cre you might ask in a nutshell it's google's realization to turn the accumulated sre knowledge public making it a de facto industry standard you should definitely check out their video now sre everyone else with cre i will link it in description it is in my opinion very enlightening in fact i think you should watch the entire series even if it's just a refresher because i'll use some of their statements to underline my following question to you is google a reputable source i say yes the wikipedia article might need a little update if policy allows and so do you think our valued partners the cloud people who are by the way servicenow and gcp certified that being said next week on june 24th alex will have a joint item master class webinar together with one of their smes covering the sre's guide to servicenow and observability from a business stakeholder perspective so do not forget to register or if you are watching this after set date find it back in our youtube playlists together with many other interesting and related topics just like this youtube video now let us quickly review the fundamental principles of sre first up reduce organizational silos for any servicenow veteran this sounds suspiciously familiar in the context of the topic at hand it comes down to sharing ownership across teams it also means using the same tooling to make sure everyone has the same view and approach when working together there is a common understanding the second principle is accepting failure as normal no system can be perfect so we need methods to mitigate this fact on the one hand we have blameless port mortems which should help us avoid the exact same failure again on the other hand we use risk and arrow budgets to define how much the system can go out of spec the third pillar revolves around implementing gradual change from a developer's point of view small incremental changes are easier to review combined with the sre practice to roll out new features to only a small percentage of the fleet it reduces the impact and does mean time to resolution short mttr while making it easy to roll back if something causes a bug these three pillars one could see them as sre process layer are underpinned and enforced by automation and tooling and the mindset to measure everything i do think you might get what i'm hinting at in a nutshell to me many of these mentioned points sound like a perfect job for service now and fit perfectly into many aspects of the platform do not get me wrong i'm not saying that servicenow is the holy grail for all of sre's needs however it comes with many relevant and attractive tooling being an enterprise service management and automation platform that is servicenow also allows for a hybrid model between traditional ideal based it and a devops driven one as a matter of fact sre relies on idle practices such as incident management and event management that is also precisely how servicenow positions sropes in its product catalog so do mind the dependencies srops promises to deliver a solution for a hybrid model addressing agile devops and sre teams that require faster onboarding and a new lightweight solution combining item health and itsm the site reliability ops workspace provides a unified user interface leveraging the new now experience ui framework in the current version quebec that is it provides features such as rapid team onboarding to create and manage teams using a flow which should make it easily extendable using integration hub and its various spokes also because we leverage the platform's incident management process meaning we can use on-call features we do also get an additional change request type that is pretty cool considering it allows for a dedicated change request workflow secondly we have rapid service registration fancy talk for letting teams create microservices in the cmdb as application services and relate them to other application services and business services i can feel how some of you might now raise an eyebrow or two i will not open that can of worms for today nevertheless as it is very tempting here my two cents application services are a starting point for service mapping which are also conveniently featured in the workspace however being a new solution a lot of functionality is still missing and by the way registering service can also be done through available apis very neat now speaking about integrations srops leverages out of the box connectors for common notification and collaboration solutions twilio teams slack most importantly however are the connectors for telemetry integrations these allow a site reliability engineer to create actionable alerts in an apm tool and run them directly against own registered services in servicenow the srobs workspace conveniently provides a webhook url for each new alert integration and of course the means to take action when alerts do arrive taking action brings also the opportunity to automate more using integration hub i can see how this plays a very prominent role in the anti-sre automation and tooling discussion but i digress now taking all of these these means together enables sre teams to monitor service health and allowing them to respond to alerts and incidents in a unified way additionally the aggregated data presents itself perfectly to measure well everything as a ropes in that sense delivers the necessary tools to define sli slow arid budget policies etc i will probably have to cover those more specifically in a dedicated video at some point that said there is much to discover and discuss i must admit writing this episode took me much longer than i expected researching the various topics and dependencies is comparable to a field full of rabbit holes down you go nevertheless if you care to join a lively discussion about a specific topic you can find me in the various servicenow communities just ping me wherever and i might send you an invitation to my corner alternatively it would already mean a lot if you hit the like button and subscribe to this channel that being said thank you so much for watching and see you next time [Music] goodbye you
https://www.youtube.com/watch?v=RyF1GQKDMlk