logo

NJP

ITOM Masterclass: Site Reliability Engineering & Observability @ ServiceNow

Import · Jun 24, 2021 · video

[Music] all right well good morning or good day depending on in which time zone you are beloved audience um super exciting that you could join here today's master class about itom and more specifically the topic around sre so site reliability engineering and how that relates to servicenow because this is a very hot topic we have seen service now than recent acquisitions in the space they they have bought up some technologies and in general i think that we're seeing the industry trend when we're speaking about iot operations especially moving towards the world of sres hence we thought it was a great idea to arrange a one-hour session where we just kind of touched the basics concepts around it and and really kind of dive into it and what it's all about especially in the context of service now um just some housekeeping things before we get started this session is being recorded so for all of you on the webinar and or for those of you who are watching this at a later stage um then yeah you will be able to find it on youtube it will also be sent out um in the email that is going out after the webinar um as well as there is a qna box which means that you can actually type questions and you can ask us in the speakers any questions that you might have and we will either read it straight in the webinar and try to answer it or if we don't have time we will get back to you personally later um so i think that is it from a housekeeping perspective we have roughly one hour with 58 minutes left but before we get started i think it makes sense to actually give a proper introduction of who we are so me who speaks right now my name is alexander youngstrom and currently i'm patching in from amsterdam where i'm located originally i'm swedish though um and it's actually a sunny day here so it's quite nice with the summer coming back but i work as the managing director for anal and partners and personally i am super passionate about this topic of sres because what we see is that the world of i.t operations is becoming more and more complex and everything that that has to do with it so machine learning people and processes that is something which i am personally very enthusiastic about so in the past i have also worked at servicenow in their itunes team and i've spent the best well of i think 10 years only focusing on on this area pretty much so super excited to have you all on board here today and speak about it luckily i am not only myself here today but i have my good friend tapia with me so tapio who are you oh thanks uh my name is tapes i am a senior solution consultant at the cloud people i've been working with servicenow for approximately eight years helping customers with their it operations and mainly uh getting them up to speed on modern practices and help them develop their business super cool and where are you located again tapio with the audience i'm actually in the center of finland in europe center of finland all right all right super cool i see we already have a question right now from stefan in the audience and he is asking will we have an access to the recording and indeed you will all have access to the recording so you can send it to all your friends um all right so what do you say tapia should we kick this off then yes please all right um now just high level here what we will go through today um of course the area of observability is huge and we're kind of opening the pandora's box here a little bit so we have to be disciplined in where we put the level of let's say knowledge it shouldn't be too high it shouldn't be too low and because it's a very broad topic um we cannot touch upon everything with that being said um the areas you see in front of you here mainly redefining the cmdb so as we all know the beloved child of cmdb has been a discussion for the best 12 20 years soon how does that actually needs to be redefined in the world of sre um we will speak a little bit about connecting data so how can we connect different data points and how can we ensure success there how can we act on data so automation machine learning that type of topics and then finally if you stay to the end i will speak very very practical from a servicenow perspective about the journey so this is not your usual sales type of lingo or talk but it's very practical advice based on over 30 implementations that i have overseen in this area of how you actually can reach um further in your journey of sre operations but before we get started i want to speak first a little bit about the complexity and madness that we are seeing in the world of i.t operations today and me personally it's all about finding this what i call golden formula um so in the terms of complexity and madness um there was recently a forester report and in this forester report there interviewed over 400 enterprises and the data that you're seeing here on the left side are the results of that forester report i have only picked out some of the kpis here but essentially 25 of the organizations were using 50 or more monitoring tools 40 of the organizations were receiving more than 1 million events on a daily basis and 11 of them over 10 million of events on a daily basis and this is just speaking about monitoring and events really because this picture in front of you here is something which is very common and that we're seeing today that in the past iit operations used to be fairly centralized it used to be fairly easy it was the it department they had their servers they had their databases they had their data centers they had their support stuff now everything is i.t everything is data you have probably a lot of you have already heard the terminology data is the new gold and um i think that although data is the new oil is what they're saying and i think that's that's very true um so we see so many different data points so many different variables anything from operate to delivered to how we build things to how we plan things employees partners strategy customers integrations clouds yada yada you name it and all of these things are generating data so really what we're trying to do um all enterprises today is the first step is really trying to take all of these data points then try to categorize them and then finally to truly structure them if enterprises are successful with this mission in my opinion that means that they are also truly on the path towards becoming data driven um without using buzzwords but but this is for me what it's about when we say being data driven really um so for sure it's very complex and i i don't know tapio i guess you have seen a similar trend in your past eight years when you have working with enterprises right definitely yeah yeah so um it's really crazy and then if we're looking then quickly kind of just setting the stage here or let's say establishing the nature and definitions of sre practices so by the way um i think it was 2003 already google introduced the the terminology sre and it stands for site reliability engineering and whenever we speak about sre practices it usually have a few things in common a few pillars and this is not all of them granted i have picked out the ones which which i find the most interesting but starting from the left bottom side here then we always see that srs tends to be decentralized so what we mean with that is you tend to move away from having a centralized i.t department to have decentralized autonomous teams and the reason for this is that each team being allowed full freedom and flexibility the philosophy at least is that it should increase productivity and kind of fail fast fix fast approach um but it's really all focused on it service reliability and of course you on the call might now think but hey isn't this what we have done in itil and it for forever for 20 years and yeah that's true but when we speak about it service reliability from an sre perspective it's really about determining what is the specific level of service reliability that we need to meet for different applications or services and what are the earliest and the best indicators let's say um for how we can indicate a service degradation so it's all about finding those those benchmarks and those kpis and and often what we see here is that sre teams they work with new concepts such as error budgets they are allowed to have failures um slos so service level objectives so for example what should be the uptime what should be the availability and that type of things so it's a little bit of a shift compared to the normal traditional itil based approach but yet the end goal in my opinion is the same which is iot service reliability really um then of course we have quantifiable monitoring so things should be measurable how do we actually monitor and ensure this service reliability how do we automate things because automation is key and utapia will speak about this just in a bit and then finally what we also see when we speak about sre are cicd and devops so these things together um constitutes the majority of sre activities i'd say and this is a huge shift for a lot of enterprises and a lot of organizations to try and incorporate this new very dynamic way of operating them oh now i went here to the too fast so with that being said speaking of automation maybe you could go ahead and share the screen from here and yes i will do that speak a little bit about toil before we jump to that i want to add one little notions here to this point uh sorry originally wasn't something that google was public about that was something that they kept as a secret inside their organization and they only later on came uh public with it so it was their internal secret or their secret sauce as they called it okay i didn't know that actually yeah uh and i i think it's uh you can sort of see it in some of the practices uh they they sort of kept something secret first okay uh going on to uh to the components here so can you see the next screen i believe you do so the thing that we're uh or our topic that we're gonna discuss is toil i'm sure many of you have not heard the term and well some of you most likely have but there are still a lot of who have not heard the term and i want to come back to this in a bit so that's why i'm going to define it uh toil is a task that is usually manual something that your sres will go and do it's usually repetitive which means that it's not done just once you will have to do it again and again obviously when there is something that is done by a human its manual process and it's repetitive you can automate it then you would actually actually ask why why not just automate it straight away and there there's the development life cycle um when you have development and then you have toil they sort of uh try to cancel or totals starts to slow down development uh toilet is off often reactive so basically you don't plan it it comes to you and you have to take care of it something happens in your environment or your your systems or your services and you have to go and fix it so you're responding to something happening then very very crucial indicator here is that it doesn't have an enduring value meaning once you go and do your action you fix or you close or complete the toil uh it doesn't mean that it won't appear again you'll most likely have to do it again at some point basically you're you're going from a change state back to the previous state that you are in so your system will stay in the same state as before and the worst part of this is the amount of oil the volume of toil grows relative to the growth of your service and your business and that is a very very crucial point here uh and we'll talk about that a bit more just one thing before you continue um some practical examples here maybe would you say that toil um could it be for example running scripts rebooting servers rolling back you know what what type of activities could be categorized as toil in your opinion uh i'll like i'd like to give an example from the source number so a servicenow often relies on workflows you define something a process and then you try to automate it when you're designing your flows you taking into consideration the state that your environment is now and perhaps some changes that you're expecting to happen but we just don't have the capacity to you know predict everything that's going to change and then the the processes and the flows that you've designed they lack a capability of doing something automatically and then something unexpected happens in your system and you have to go and fix it oftentimes you'll create a script for for for example uh from my past experience uh we had a flow for change change management and it was incapable of uh of switching uh there's a specific state if something occurred which required you to go and fix it by a script now that is fairly simple after the first time you'll have to go and do it again there is a very small part of human validation there but technically the evaluation could be done by a machine as well so then some operator goes and runs a script to fix the state back to what it was before so that your process continues on basically something that you need to do to keep the wheels turning keep the lights on that kind of thing right okay yeah thanks a lot that makes sense cool yeah and that actually uh describes the first uh con so it's an unnecessary necessary you have to do it uh enable to in order to keep your wheels turning and lights on and the biggest con here it causes latency or delays your service and your development development being the most crucial part there even if there is a short delay in your service that might be acceptable but in the long term latency and slowness in your development will be very detrimental to your business because in the modern world you have to compete with all your competitors and the best way to compete is by constantly developing improving your business improving your services and toil is definitely one of those things that will hinder your ability to develop and then the third con is it is most often really good at causing stress for your sres obviously there are different types of people some people are perfectly fine nine to five just doing toil they're they they're happy with that and that's all they they want but then i would say i would argue that most people are very unhappy with toil me going into that category and yeah toil is something that will hinder not just your development but it will have long-term effects in your business because it is causing stress and hot and happiness on in your employees in your development staff it'll it'll cascade it to other other events meaning your developers might be unhappy and leave the company trying to hire new people might not be that easy since uh the reputation of having to do a lot of toil can be something that people do not want to you know be associated with or don't want to work with now there are some pros to toil usually where where things like this happen or why there's or what toil is uh associated with or what causes toil is something that is related to your core business and you you can easily use these tasks toil tasks as an on-boarding task for your newcomers so that while they're fixing the the state the situation they get an overview of our or an understanding of the system they see what what the flaws are they they have to look under the hood to see what's there and it provides them some visibility to your your business and technically uh also it can give you instant gratification basically there's something wrong and once you fix it you feel like okay i accomplished something today normally you might be in a meeting for eight hours a day and you feel like you didn't accomplish anything and then then you get a task that requires you to reboot a server or something and then you go and do it and feel that i actually accomplished something meaningful today not just discussion or not just something that i didn't find value in since it does provide a very small bit of value as your services up and running so could you say tapio that it's an illusion of value yeah i i kind of like sugar you shouldn't really eat it but you still wanna and you get a little high out of it afterwards but yeah unhealthy in the long term and then one pro that you could save for it is is it can offer you a break in a hectic development environment i am 100 sure that a lot of you have most of you have experienced the hectic world of development where you have deadlines and you have friends and whatever and you're you're constantly trying to reach your goal and you're looking at the calendar and you see that they come in there and then when you reach that uh date and you publish and then you might afterwards you know being you know processing toil and then you sort of get a breather where you're no longer so much tied to deadlines and and schedules so that is a brief brief explanation of title uh i'll go into the next point here which is managing toil now google provides before we jump in here and that from a timing perspective yes you're aware we have roughly 40 minutes left just so you know yeah i'll be quick i'm i tend to talk slow a lot okay google's point of view is that you should aim for a a maximum of fifty percent of total per sre how google tries to accomplish that is by having weekly rotations on their operators or developers that you will do one week of primary on-call where you just handle oil secondary next week is secondary on-call where you're also doing toil and then you have two weeks without it and that way you'll get 50 so you can really easy to calculate there you need four people for that minimum four people or maybe with two if you just do one week total one week development that really doesn't work in the real world but if you have a larger team you're able to you know spread out the weeks you can have one week primary one secondary and then more rates of development and that way the maximum amount of toll that you can have gets smaller uh google's own survey says that their employees have about 33 oil which i still think is kind of high the same yeah yeah i i wouldn't be happy with 30 work a third of my work time being put into just keeping the lights on so that's why i would say that you should definitely aim for less than 25 which means uh a quarter of your time or per well you would like to have it less uh and if you have or if you don't have the capability to do this rotation if your team is not big enough if you maybe have only one person uh doing this management then you are forced into the situation where you definitely need to manage how the toilet occurs you have to focus on automating it before you expand your services because very often what happens is your sales is pushing your your business to grow and obviously you have a push from themselves which is saying that we gotta onboard new customers we gotta expand our services and at the same time development operations are are saying that we have so much toll that we just can't do it and then you get to the argument okay let's fix it after we onboard this new customer or provide this new service but like i mentioned in the previous slide the amount of toil is relative to the the size of your service or as your service grows so you might be setting yourself a trap by not taking care of your toilet sources when you still can i love that you say that because i think what we're often what people often think is that the problem will be sold if we just hire more staff or throw more support engineers on it or something but clearly it's it's a matter of addressing the root cause is what you're saying yes and with sre's ideology is measure everything once you understand what coil is you need to measure you need to understand how much we actually have without understanding what it is it's really difficult to measure you might because again there are some manual tests that are not spoiled if there are some stuff that you got to do that actually brings you value or you don't do it again another time so import is to measure how much foil there is plan on how to getting rid of it so not just saying that we gotta get rid of all of it and that's that's it plan it and then eliminate it you might not be able to remove toil completely perhaps you can just automate part of it and reduce the amount that you have to spend the amount of time that you have to spend in dealing with it so move gradually i'm skeptic in that sense that the world is never complete our services are never going to be perfect so move one step at a time that's the planning part and then the goal is try to eliminate it sadly we'll never reach now thanks a lot yeah i think this is super important when we speak about sres and a great kind of setting the stage here um because let's now um go i'm going to hijack the screen sharing from utapio thank you and let's zoom in here a little bit then on [Music] this service now specific things so the first thing i which i find important here by the way i'm speaking now first about redefining the cmdb so if we're just looking at the nature of a cmdb what you're seeing here is kind of the transgression of how infrastructure has developed over time so 15 years ago that's when we saw virtual machines you know virtual machine containers farms etc but now what we're seeing as illustrated here in front of you is way more like docker kubernetes containerized environments kind of stateless applications and infrastructure as code and that type of operations right so the fundamental issue is that originally the cmdb was created for a world which existed 10-15 years ago meanwhile what we're seeing is that people try to fit in this new world into the concept of a cmdb and that's where the friction gets created because there are so many dependencies and what i always ask myself is do we really care about infrastructure these days i mean yes we do but at the same time not always because often what we're seeing here and like tapio mentioned when a service we have some sort of service degradation or something like that um especially in a devops world it is expected that the service also auto scales so for example we add more pods if it's kubernetes and we add more containers if it's darker we add maybe more more virtual machines to a cluster automatically you name it um but fundamentally things are auto scaling and the infrastructure becomes less of importance and in matter of fact infrastructure is becoming less expensive as well hell like azure gives you i think 300 bucks for free just to play around you can create an account today um so the auto scaling is huge tags and context um if we don't care about the machines and components in the same way anymore what we have seen is that tags and context is way more important so tagging the resources what are they used for in which context are they used is it a dev environment is it a prod environment and so forth and so on um things tend to be more modular so instead of you you all know this kind of monolithic applications perhaps versus the microservice based approach but more specifically what we see is that rather than having one big application or a very simple application stack things are much much more modular now hence way more dependencies but perhaps the most critical element here is that there is an entirely different level of self-service that people expect if i am a developer in an organization i also expect to be able to very quickly create my own resources without having to wait for approval and a virtual machine and having it imaged or whatever i should just be able to spin up something fairly fast so these things together is what adds to to sre operations and if we then look at the context compared to how the cmdb was created this obviously creates a conflict so here's a bit of a controversial statement perhaps um but should we actually maybe stop calling it a cmdb is something i often question um because when i hear cmdb and when i speak to you know people recent grants from universities that are doing devops like devops engineers sre engineers or something and you mentioned the word cmdb they think oh that is an old concept who is this dinosaur speaking about a cmdb because no one wants to be stuck in the old world right so in my opinion a cmdb that should rather be called a dynamic data layer um and this is exactly what what servicenow is doing and that's why also we have seen service now created something called for example service graph so i think slowly we are moving towards an industry trend where the purpose of a cmdb being keeping track of things it remains the same but how we name it and how we interact with the cmdb will be different um so in front of you here you have three areas and in my opinion then a modern cmdb should support all of these areas so first of all it should be component focused so be able to capture the traditional i.t because whenever we like it or not then we still have data centers we still have network equipment virtual machines we cannot just ignore those parts especially not if you're a big enterprise traditional i.t will continue to be a part of our reality for many more decades i'm sure of that yet not everything is traditional i.t so we also need to stop being so focused on components and rather be a bit more service focused more specifically the cmdb should be able to support cloud resources and stateless things and microservices this is really important in my opinion so not only having that component focused but really seeing like api gateways microservices infrastructure as code functions those type of granular things and then last but not least state focused so a cmdb should be able to ingest metadata and telemetry and you might be wondering hey what are you meaning alex with metadata but things such as tags things such as who owns a resource where is the resource build what are the cost center that type of metadata but through all of these layers there should be relationships and contexts the picture in front of you is for me what constitutes a modern cmdb and and this should be at least the minimum level of expectations when you invest into a cmdb that it should have the ability to support all of these spectrums um i really like to highlight the the very bottom part there the relationships and context uh even if you have all the data if you don't have the relations between the data the data is unusable usually yeah yeah exactly exactly absolutely true absolutely true um so we're really speaking about connecting the data and dots and if we're going to fill that up then you have something to share here tapio i believe about regarding the discovery part right yes let me share my screen okay uh the next part was the populating of the cmdb i split this up into two categories the traditional discovery and cloud discovery traditional discovery as i let you mention that it is for the traditional hardware servers the switches and so on servicenow itself provides tools for that there's servicenow discovery and also it supports if you want to do some other third-party discovery tools and then just populate populate the cmdb and service now with it there are agentless approaches where you just blindly look for things and based on what you find you look a bit deeper into it and gather information that way then there's obviously the agent-based solutions where you actually have an agent on the device itself and that will report back to you i'll usually consist or has schedules so you're telling your discovery one-on-one every every day at one o'clock at night or weekly or monthly or whatever and the traditional traditional infrastructure usually has a longer lifespan a longer lifecycle servers aren't usually set up for just a day uh they usually live a bit longer obviously because the cost is a lot higher yeah and back in the days that was the big thing it was that was the main information or the main reason you needed information was the cost of acquiring this infrastructure um then the the modern or the the cloud-based discovery basically it doesn't rely on the old anymore and can't rely on it usually api based instead of sending blindly pros to it you ask your cloud provider for example provide me the data that you already have in your system so it's kind of like outsourcing discovery to your cloud provider uh as you mentioned tag based tags provide you crucial information about your system which can cut down on your costs obviously which we want to do drastically it can be event driven so that instead of reacting based on the information we you know might get daily which might already be old we can move switch it to uh the uh other point of view where when things happen we alert the system about it we created events and usually the thing is that the cloud-based uh resources are very much shorter in life cycle like containers and so forth they might just exist for like 10 minutes right exactly and that's why that's why the traditional discovery just didn't cut it here since if you're running daily the state has has changed definitely has changed hundreds of times during the the time between the discoveries so you're not actually getting good data you might be getting just completely random stuff and that's why um it has to go the other way around when things happen you learn the system but the issue with both of these discoveries usually for businesses is businesses haven't defined what is valuable to them uh from the traditional side some businesses from discovery populate everything from their disks to this partitions and everything and they might not do anything with the data the same thing with the cloud how if your business doesn't use the information how many pods you have it's completely a waste of uh resources to figure out how many you have or store it somewhere or or create events or anything yeah you need to actually use the information that you get from your discovery services and to try all this uh service now or yeah servicenow has a tool for creating your relations traditional discoveries just create the finds the things you have to link them you have to link them together to get the real value out of it but you also have to manage the data you have to be sure that the data you have is accurate if you have 15 million data points and 90 of that is inaccurate then pretty much all of it is in that coin yeah and and to add on to that so i'm just going to take the screen sharing from here because what i always hear is that when we're speaking about the discovery in terms of especially docker and kubernetes i don't know if you've ever had this question tapio but like i i've often get the question of alex do we really need to know when every pod is created do we really need to know every container like what we really care about is our services and what i mean with services is like microservices right so this for me is like a crucial element when we speak about the cmdbn this is actually one of the greatest challenges of a modern cmdb so i'm going to try and explain this in a short way here to the left side you see a few different examples of clouds so we have here azure we have aws and we have google cloud and what happens is that in each cloud people have their devops things like kubernetes or they create some sort of app service or like these functions that i spoke about before basically microservices and apis the problem here though is that each cloud becomes indirectly its own little silo and in average enterprises today use at least five cloud providers and if each cloud provider have their own little world then it's very difficult to see how things span across the domains here so when we speak about microservices it's often a question on how to model that like yes we know that there might be 20 docker containers running somewhere but we don't care so much about the containers we know it's in azure somewhere we know it's a kubernetes cluster somewhere what we really care about is the micro services like the apis that humans have defined the code that humans have written the developers that they have written so how can we model this in a good way well this is all about topologies um so more and more at least in customer conversations that i have we see that the devops teams within enterprises are struggling with this they have their repositories like git or git lab or whatever it is but very quickly you notice just like how complicated it becomes to keep track of microservices because what we often see is that like a microservice rely on another microservice which rely on another microservice etc etc etc um so how can we register and how can we keep track of the microservices um well in service now there are two ways of doing this one is we can actually discover microservices and we do this based on tags so practically what that means is that if you have an existing tagging structure let's say in kubernetes in azure or something like that then servicenow discovery like tapio mentioned can actually grab those tags and based on those tags we can create service maps out of the microservices so we don't necessarily need to know the docker containers we don't necessarily need to know what is the infrastructure behind all we need to know is that this piece of functionality has this tag we should add it into a dependency view basically um and this is what is being known as tag based service mapping i think that a lot of things are moving towards that world with tag-based service mapping now needless to say not every organization have like a super strict governance and and keeping super strict track of all of their tagging operations so what we also can do in service now is there is a feature called sro or site reliability operations and developers can then actually register their microservices in servicenow i'm not gonna go through in detail today about this piece of functionality but this functionality you should know has an api and i have seen very many examples where essentially every time you create a new microservice you know cicd devops type of thing it gets automatically registered and added to the cmdb um so these two things are possible to do today it's a little bit tedious the the journey to get there but it's absolutely worth it especially if if devops and microservices is starting to become like a fundamental part of your i.t operations um now the role of components then because i think it's always very very easy to sit here and bash you know the old world and say that all the traditional i.t is outdated we don't need to care about it anymore but that's not true traditional i.t is still a huge part and the role of components it's easy to say oh don't care about it let's only care about cloud discovery and microservices and devops and sres i absolutely agree we should put more emphasis there but yet the role of components virtual machines to traditional infrastructure it shouldn't be neglected and i always say this and i just like to mention it specifically like this now i spoke before when i showed the cmdb about the telemetry and metadata and really telemetry and what i mean with that is monitoring and whenever i hear monitoring or when i speak to people about monitoring they are typically imagining these traditional tools of monitoring where you define your thresholds maybe you have an integration to service now so you create an incident when there is an alert but it's all kind of reactive and you create the thresholds as you go and that type of things so if we now zoom out the picture a little bit and we speak about site reliability engineering and monitoring it's actually a topic which is called observability and this is something that you will hear a lot about um because a lot of organizations they want to move towards observability um and if you ask 10 different people you will get 10 different definitions of what observability means so this is my interpretation of it but i like to make it simple um first of all something like this is probably what you're used to when you know we start to explain or people start to explain how their architecture looks like as i said it's hugely complex and and it's just you know you are taken as a mad scientist when you try to explain the architecture um but observability and monitoring then it's more than just alerts and thresholds if we truly speak about observability in a serious way then we want to include at least these four types of data points so we speak here about metrics we speak about logs and analytics we speak about the normal traditional monitoring data and we speak about topologies and the cmdb so these four data points um are extremely important if we want to to stay on top of observability and by the way for those of you who might wonder what what does alex mean with observability i'm i'm i mean in order to observe the state of a particular application the true as is state how that application is doing right of this minute that is what i mean with observability and then you need many data points these days but the thing here is that these data points on their own only tells part of the truth but together and combined is when you really can start getting you know observability on steroids pretty much and that's really where servicenow comes into the picture with their machine learning capabilities then and of course servicenow is not unique with this you will find other solutions out there like dynatrace and mood soft and pretty much any aiops tools today you will be using machine learning um but this is one of the few areas where machine learning is actually working and doing a good job because a few practical examples here when you apply machine learning here you can do things such as correlations and groupings you can start predicting alerts you can start doing anomaly detection you can start assisting with root cause analysis why well simply because there is no way a human can stay on top of all of this data and read it manually remember what i said in the beginning that i think it was like 50 or organizations or 40 percent generate 1 million events per day or more obviously you would need an army of people just sitting and looking at those events so this is a good example where machine learning realistically can be used but machine learning on its own is worthless um we need the people in order for it to work so more importantly um we often hear need people that can actually look at the results of machine learning determine whenever it's you know it's good or bad because machine learning on it on its own it cannot say whenever something is good or bad right or wrong negative or positive most of the time what it can do is tell you if something is abnormal or if it is normal you still need people to be able to give that human perspective on things um but also here when i speak about people i'm also referring to uh like the culture of sharing data between teams uh the culture of thinking less in a domain-centric view but more in a holistic view between teams for example processes that type of thing so a lot of talking a long monologue here for me but this is for me what what observability means alex can i add or actually a question to you yes um well we have observability we figured out what the tools for it are how we can utilize machine learning for it and then provide information you know pre-processed information people who make the decision or who will then then see the process the information regarding observability what's the function of observability why do we need it what do we do with it for me observability is all about you know we've heard before we shouldn't be reactive we should be proactive but for me observability takes it even a step further that we shouldn't be reactive and being proactive it's not enough either we should be predictive and if we can observe how things are operating under normal conditions then we can also predict things when they are not operating during normal conditions let's say so that is for me at least how i view on observability personally i agree and and usually these things are considered as something technical something that you're monitoring or preserving your services but there is a business side to it as well because your business will need to make the decisions of what we're going to do in the future based on where we are now and where we where we were before yeah and that's information for sure and they will use the data of of i.t systems to base their decisions on that's that's for sure exactly yeah no absolutely um so we have this observability maturity model um now granted i'm not gonna have time to go through this entire observation observability maturity model now so um you will get this as a sent out after the webinar it's a pdf you'll get it just as a nice giveaway for free where you can kind of see how it looks from a servicenow perspective and to move from having this kind of reactive human intervention to machine learning assisted and finally to a predictive and observability stage okey-dokey with only 10 minutes left i see that we've had a few questions a really good questions from ilka here i am not sure if we will be able to take those questions straight here now on the call with only 10 minutes left but what i'll do is we'll reach out to ilka afterwards because this is some some quite interesting questions and maybe we will create some sort of follow-up with all of these questions so we share them with the audience um but the journey then um we've been speaking about conceptual things we have been speaking about toil um we have been speaking about the nature of sres kind of how things have developed and what direction it's moving towards now let's take it down to earth um how can we really get here using the servicenow platform well first of all there are some organizational pillars that needs to be in place um so from a cultural aspect and you and me tap you just before the call here today or the webinar had some some good interesting discussions here um for me it's at least about sharing data um sharing data between the teams knowing kind of how can other teams access access my data what data are we generating if we're speaking about observability how does that data look in terms of monitoring in terms of measuring you know how can we truly share our data and create that sort of open embracing data culture within the company that for me is super important i guess you agree there as well to help you yeah definitely and not just about sharing the data but perhaps even uh just giving an opportunity for your other systems to get access to the data if you don't want to push it everywhere all the time it can be in the future your other business or other services might want to have uh some of the data that you weren't you weren't sharing before but yeah having access to it and allowing all the other services to expand to with knowledge from other services as well that's right yeah exactly exactly good good addition um and then of course as i mentioned in the beginning if we are serious about sre operations then considering reliability super early is very very important so whenever we're provisioning a new service of some sort whenever we're considering to expand a service introduce new microservices whatever it might be reliability should be included like from day one you know we do not launch a service at all without having very very in a minuscule level um consider the reliability parts of it um and finally i think it's super important to try and evolve beyond itil i am not saying that we should throw out itil through the window because incidents and changes and so forth will still be a valid concept but at the same time we cannot be stuck in this in this bubble of itil all the time either but we need to grant some flexibility there um and of course that's almost a philosophical discussion and we could probably have debates around this but that is at least my opinion from an organizational perspective so these are the main organizational pillars and the cultural aspects i would say if we're now looking here as one of the final slides on the technical pillars then to reach observability and i'm now speaking from a service now platform perspective and i am first assuming here that a basic itsm is set up and by the way these things are based on reality i have worked as i said before with like we have worked here together uh with like 20 30 customer cases um in these areas so the first thing is i think it's a quick win to integrate any cicd pipelines so devops operations to change management there are huge benefits with that simply because we want to expose the world of devops to the world of itsm and iit operations we don't want to let the devops you know run wild on their own necessarily but we want to give that transparency and visibility in what is actually going on it doesn't mean that every change needs to be approved by the way but it should be integrated in in some sort of way so that's the quick win and you can do this without any additional licensing model or anything like that and then i always say that creating that data layer if you haven't started that work already so aka acmdb hugely important and i would prioritize here the visibility to cloud if you have that the reason here is simple because if you have visibility to cloud if you know what resources have been created what has been decommissioned what has been spin up over time etc trust me the bills and the invoices you get from amazon and assure google will make way way more sense so getting that visibility into the cloud is is kind of a key here in my opinion but then of course as well as i said don't forget about the the on-prem and traditional i t then connecting basic telemetry it doesn't mean that you need to be the very best with the latest dynatrace and apm and all of these things but if you have some existing monitoring tools connect them to servicenow um if you do these three things it usually takes six to 12 months and trust me your organization will be on a higher maturity than i would say 80 percent of the most common ones um if you manage to to build these three layers here then you are very very well far off um into the maturity scale then you need to ask yourself and consider what is truly important for us on a long term basis um are we interested in these service topologies like microservices like i mentioned tag based service mapping whatever it is are we more interested maybe in aiops capabilities like machine learning um or is it more like the automation parts that that we truly are interested in the reason that you should ask yourself this is that depending on where your priorities lie then you might essentially be able to optimize the licensing models and the cost and implementation and so forth but if you do these things then you really are only let's say cutting edge then then you are pushing the servicenow platforms to its limits and you know still within the capabilities of what it can do but then you are ahead of a lot of other organizations um so these for me are the the technical pillars so to speak um and of course there are other aspects to this i don't know tapio if if you'd like to add something particular but i think the basics here of the quick wins the data layers and essentially connecting events is more than enough for most organizations at least yeah definitely and in addition to your the considerations part when you're assessing if your business is interested in these uh if you say no then list why so because usually your decisions are based on the current situation situation always changes that's the only constant yeah yeah absolutely absolutely um so with that being said um this was really a scratching the the surface here a little bit um of what observability means the decide reliability engineering practices and how it relates to servicenow and granted we have only scratched the surface we have only spoke about it on a high level perspective but if you would like to have more conversations about this you can scan the qr code in front of you and you will be connected to either me or tapio there will be a follow-up recording on this and of course you should follow the cloud people or internet partners on linkedin just to stay up to date and in general this is an area which we are very very passionate about and we would love to to hear more from you if you have any questions or if you have any curiosities in it but with that being said i think we are within the timeline limits tapio and yeah at least from my perspective i'd like to thank the audience very much for their their time here today so what do you say should we wrap it up yes and thank you to all the audience and have a great midsummers yes exactly have a great mid-summer guys super cool thank you [Music] all you

View original source

https://www.youtube.com/watch?v=XQk2D9-AhAs