logo

NJP

Masterclass: Demystifying ML- & AI in ServiceNow ITOM (Metric Intelligence)

Einar & Partners · Dec 09, 2021 · video

[Music] because we have a pretty packed schedule right okay so we have a lot to cover [Laughter] so let me go ahead and start sharing my screen here um okay you can see my screen and keep yes i see you perfect perfect so excellent um well a very sincere welcome to everyone who have patched in here to this master class today and it's as always super exciting to see so many people showing interest on linkedin so many people signing up for these types of master classes it really makes us motivated to continue making them and to provide this sort of sort of knowledge for for the professionals out there and today's masterclass then that you have logged into is about ai in servicenow item so well this topic has always been a little bit let's say mysterious because a lot of people when you hear the concept of ai then there tends to be a lot of assumptions of what ai is and what it should do and there is essentially a lot of confusion um so me together with my colleague aquifer today we sat down and we had a discussion about can we try to demystify this concept of ai especially in the context of servicenow item then and that's what we're going to do here together today um but before we get into the actual webinar then i think a little introduction makes sense so me who speaks right now my name is alexander youngstrom and i am the managing director at ainan partners for those of you who have previously attended master classes by inam partners or who have followed us on linkedin you probably have seen my name once or twice or thrice maybe but i am very very passionate about it operations and i'm also very passionate about the topic of aiops or artificial intelligence operations so previously i used to work at servicenow in my beginning of my career in their itunes team and since then i've had a long long journey when it comes to iot operations but these days really what i'm doing is that i'm advising clients and companies around the world on the strategy when it comes to servicenow item however i am not alone here today but i have a good colleague with me a keith so um who are you akief and what is your background to this field of ai because you have quite a good background so yeah so thanks alex uh yeah and uh akif i'm the r d lead ai ops and data science at inarm partners and um yeah i have a background in econometrics data science ai and uh yeah quantum computing and i i what i love to do uh it's you know all making sense of data like modeling and uh creating statistical models and um and yeah like uh demystifying actually also um you know ai applied anywhere actually and that's what i love to do and and i'm really excited to uh organize this master class together with you yeah excellent super cool to have you on board and we're gonna have a lot of fun here so um now for the audience here we pretty much have 55 minutes and this is a huge topic when it comes to ai and it's absolutely an enormous topic if you look outside of servicenow so we will have to try and time box it a little bit but kind of just to set the expectations um so what you can expect to learn today first of all when me and a key if we sat down and we planned this master class we thought it makes sense to just kind of define the very core concepts of machine learning naturally this has been done before but i think for a lot of people on the call here today and it's kind of difficult to grasp what really is machine learning what is ai so we want to set those concepts straight from the start here then we will speak a little bit about how we can apply the topic of ai the topic of machine learning and especially them in the relation to metrics and models and classifications um i know what some of you might be thinking now that well alex this sounds very very technical um it's a little bit technical i'm not gonna lie because ai and machine learning is a technical field but we will nonetheless do the best to kind of give um let's say a high-level overview on these topics so for those of you who think that you will be sitting and learning algorithms today that's not the case we might have follow-up webinars about that but today we're scratching the surface here um but an important part when we speak about machine learning when we speak about ai is predictions and you actually if you have done some good work here of kind of debunking what predictions really are so i'm looking forward to hearing your take on predictions uh because as we all know predictions is a big thing in iit operations in monitoring and so forth um actually talking about the future so especially talking about the future exactly exactly um but more importantly we want to try and bridge all of these theoretical concepts to how that can relate to the journey of the service now and kind of the the relationships between all of these different concepts then so i hope that makes sense um and just some practical information before we continue so we will have a q a session in the end which means that you should see one of these q and a chat functions in the in the zoom webinar which means that you can then write to me or a keith and hopefully we can get around to answering the question if we don't have time we'll reach out to you personally with a reply additionally this is being recorded which means that if you have colleagues if you have friends etc who would be interested in this but they are not here today you can send them the recording afterwards and of course we will send you the link etc to that as well so with that being said uh what do you say okay should we get this started yeah let's start yeah and i think you are the first one here because we had a lot of interesting discussions about this phenomena of intelligence so why don't you take it away here and let's dive in deep definitely and let's let's dive into the intelligence part of the ai ops and intelligence it's it's a phenomena i mean uh you can't um see it um but this is not the really like tangible product of it it's like a force you know um you see only you observe force when you act upon it and there's some intelligence process of the phenomena and it starts all with the data and as here in the pyramids in the bottom um let's go through the steps of the intelligence process as a process and so when we have data and we structure and pre-process and data we get information and once we have information we can start modeling it and analyzing it and once we analyze it we get knowledge we start to you know understand it you know what's um uh what is the data all about you know and intelligence itself is all about making sense of data and once we start interpreting the knowledge information and data we kind of get the you know phenomena called intelligence and once we have a good understanding knowledge and um in a sentence make sense of our data we can sort you know act make decisions upon it and act upon it and that's actually the top part of this pyramid and all this process uh if you put it in the context of ai ops it operations um we create we can create this artificially on machines so um intelligence is a phenomena and um yeah we talked about ai and you know how can we do that how is that actually created on machines so you may heard of artificial intelligence and machine learning and i want to take a deep dive into a bit ai concept and machine learning because especially later on we will talk a lot about it especially how it's applied within surface now so what do we mean when we speak about ai and ml so ai is kind of like intelligence demonstrated by a machine it's like all encapsulating a concept it's a program that can mimic human behavior and sense adapt to its environment but as you have you know observed before uh in the pyramid uh there's a you know you can't have ai and intelligence without learning without learning from the data so we have machine learning that learns and improves from experience and machine learning is part of ai it's a kind of subset of it can we say so how does machine learn there are several ways and approaches we call it it can learn um you know by supervising the algorithm so kind of guided learning algorithm that you know match the marked input uh to desired output and there's kind of like a teacher you know uh helping to learn the algorithm uh we can do it in an unsupervised way and you know we look at the data we try to discover like um patterns within the data we can do that uh maybe in a reinforced way actually much more in a dynamic way like there's an agent um learning with reward and punishment so when it learns well it gets a reward and when it didn't learn well it's a it's a punishment and there's an also a concept called the learning cult deep learning which tries to imitate actually human learning it tries to you know mimic a human behavior and scientists even try to mimic how child or kids learn and then try to put that in an algorithm so a keith a question here because i we you and me had quite a few discussions about this and if if we look at the so first of all how i interpret this image is that when we speak about aei that's really the wrapper so to speak around all of these different concepts um like machine learning supervised unsupervised deep learning etc so ai is is more the generic name of all of these different methodologies would that be fair to say that it's kind of a wrapper around it exactly uh it's exactly like so um it's kind of like that wrapper and it's a little bit more like uh you know on the top of that pyramid you know ai can you know it's also about decision making and act upon it well from what you have learned exactly yeah and i think you hit the nail on the handle because that is often but what i see the holy grail is when we speak about it operations when we speak about aiops i've heard this expression once which is ai is easy operations is hard so basically interpreting the data and getting these algorithms to work which we're going to speak about more later um that's the easy part but then acting on that data that's the more difficult part so so that pyramid i think the longer like the higher up you want to get in that pyramid um the more cutting-edge and difficult it is perhaps or more challenging of course of course and and we will get later on in uh also about why it's difficult yeah exactly and then just one more thing here because you know um you are clearly more in-depth expert than i am in this topic but we had a discussion about supervised versus unsupervised machine learning and if i understood you correctly when by the way we have both of these concepts in servicenow supervised and unsupervised but to supervise we basically tell the algorithms to look for certain patterns to look for certain things meanwhile unsupervised if i understand you correctly then it does that by itself correct exactly i couldn't describe it better and that's what exactly the um supervised and unsurprised machine learning algorithms within servicenow does so when it gets its data from the applications you know it looks um either to certain patterns that you already said you know look for that or it look at data in an unsupervised way you know what kind of patterns are there so yeah like like you said okay but that makes sense that makes sense but let's take this then into a context of servicenow item then because i think this is a good bridge then so how can we look on this on servicenow item so i'm going to steal here the screen sharing back from you and then so just like a cliff explained we hand all of these different machine learning concepts in servicenow item but i think now what i would like to try and do because now we established two pretty fundamental core concepts one is the pyramid of intelligence that we need to to try and define what intelligence really is and as we can see just interpreting data that's simply just half of the picture but you also need to act on the data and then furthermore when we speak about artificial intelligence um it's primarily within aei that we have various ways of machine learning statistical models etc so in front of you here is a maturity model observe is now item especially when it comes to the disability and health parts i'm not gonna go through the entire maturity model today but you know to make it simple from the left side we have the first steps which is always like itsm um incident problem change you know the normal stuff then organizations and customers they tend to start with the consolidation of data so they start building this data lake like we discover it with a cmdb and then slowly customers and clients and organizations they continue this journey insert is now item when it comes to service mapping event management consolidating events etc but then on the very far right side that's where we have the ai driven parts and this is where it gets a little bit confusing for people so the truth is that machine learning it exists in almost every item module in servicenow so like when i started working with servicenow item which is almost 10 years ago soon oh my god then machine learning wasn't really a thing yet but if we look today at the usual suspects so these are the modules that a lot of organizations are using when it comes to iphone in things like discovery we have application fingerprinting in service mapping we can analyze traffic and we can get connection suggestions so based on tcp connections we can get a lot of suggestions for how we're going to build our service maps and then when it comes to the event management part we have things such as natural language processing so we recognize texts and patterns in text on incoming alerts and we can also cluster these alerts so we can group them in different ways in different logical orders and all of this is done to a large extent through one method or another with machine learning but here's the kicker that this is merely features just because you are using these features doesn't really mean that you as an organization um are working with like an operating model that that uses machine learning so that's really where the aiops and metric intelligence part come into the picture so on the maturity model that was the one to the very far right and this for me creates a fundamental difference because this is truly a new way of working for a lot of organizations and only five years ago this was very very difficult to work with like aiops machine learning metric intelligence etc but these days we see things like in servicenow we have off the shelf commercially ready products which is easy to get started with but nonetheless it has like a very very big impact on the operating model and how monitoring teams are working and that's really what we're going to speak about now for the coming 40 minutes so we will not really touch upon the the machine learning parts of like discovery service mapping etc but we will really look at how can an organization take that next step and really proudly claim that they are working with machine learning and ai so that's what we're going to get looking at here now the very first thing this starts with if we speak about machine learning in servicenow and if we speak about aiops is something called time series and statistical modeling um so you and me okay we will have to try and do a good job here at keeping this like to to a fairly high level um but i want to start off with the time series data so most of you probably have heard about it before some of you maybe haven't but the time series data is like the fundamental core concept of any machine learning in servicenow when we speak about aiops so i'm just gonna explain quickly what is time series data and how does the format look and what are some practical examples and then after that you active you will explain a little bit more around the classification and statistical models yes um all right so what is the machine sorry what is time series data in a nutshell time series data is a sequence of data points that occur in a successive order in a particular time interval um maybe that doesn't tell anything to you so i'm gonna put that into a more tangible graspable concept so here is a box of some examples where we have like cpu we have memory we might have number of transactions what bandwidth is in use disk whatever it might be response times on websites etc etc the common thing with all of these these metrics that i'm mentioning here is that they are happening every minute or very very frequently and what i said there metrics is perhaps the most important part so really what we are doing is we are looking at metrics and we are saving those metrics every minute in what is known as a time series database so servicenow when you get started with ai ops then you will actually get assigned your own time series database which is a separate database outside of the service now instance it's called clotho the database but the purpose of this database is essentially to store time serious data and when you look at for example application monitoring tools and scom dyna trace etc they are you know they are reading metrics but the difference here is that these metrics are being saved in a time series format in a time series database so when we can save these time series when we can save all of these metrics in a database that's when the real interesting things can start happening because then we can see over time how are certain metrics for certain applications performing and based on that we can also see behaviors we can see trends we can see a lot of things [Music] so i think i'm going to leave it at that but i just wanted to mention here that the time series data is pretty much like the core concept when we speak about ai and ml in servicenow aiops would you like to add anything here okay or does it make sense it makes all very good sense and especially like the time dimension it's important that makes it the time series data and time it's very important dimension of it so yeah exactly and actually it's good that you say that because when we look at traditional metrics and monitoring like if you're lucky they save it for maybe a day you know you can see metrics if you visit scom or you know one of if you go to your own cpu on your computer and you open the task manager you maybe can see how has the past minute been but after that minute the data isn't really saved anywhere but that's the difference here we are actually storing that data sometimes for days sometimes for weeks and sometimes for months um yes so thanks for adding that um but then aki how do we classify this yeah we have a really nice machine learning algorithms for that and in case of like service now item we have fancy supervised machine learning algorithms like the k naive bias decision tree and maybe in a later masterclass we should go more in the technical details of it but uh to give you an impression of how machine learning is in when it's in action um like so when when the matrix data is collected um like you said uh from the applications it's fed into these supervised machine learning algorithms and then this is done in the training phase so there's like a training phase a test phase and then it can be if it has good predictive power the models can be used and in the training phase um it trains on the incoming and data um and then it builds a model and then the best model that gives good prediction against on the test data it's chosen and then used further and let's dive a bit more into this process so to make sense so um why why do we need to um know more about the data or why should we actually um uh thread it into supervised machine learning algorithms um and the reason is actually so when the operational metrics data is collected it has certain you know behavior patterns and we want to know about that and the reason is why we want to know about that very intimately or very good it's we want to create accurate statistical models based on that uh like when when we get um you know operational metrics data for example from your um asia um that's collected in in case of item from with an agent client collector it's kind of scraper it collects all these data in a time series format so it's like a pre-processed kind of and structured and then it's fed into these supervised machine learning algorithms and then this supervised machine algorithms learns about the data and then it says it contains like a seasonal pattern it contains you know trends and once we know about that then we can create like statistical models you know based on that you know because and the reason why i'll show you that in the next uh you know in the next slide uh for example in the top um uh plot uh we see request uh rate right we see a pattern that's recurring like um you know through the day it's uh high in the evenings it goes down but we see recurring pattern and on the bottom we see uh like a trendy data or trendy uh pattern and we see that it starts from monday until friday uh you know it goes higher and higher the used memory or so and then on friday evening it you know goes down and then the less memory is used and then on monday it starts again so a recurring trend and why should we know about this and so if we don't know about these patterns then and we create random thresholds for our statistical models then we get inaccurate predictions but then you know why do we create statistical models and why are they important well statistical model like um model kind of the data generating process you try to kind of model actually how the data all these time series data operational data metrics is modeled and why do we want to do that because we can then make predictions we can you know make statements about the future so um and that's you know i i that's a part i love from like uh data science and ai ml and uh econometrics so um you know talking about the future so once we capture uh how the data is created once we model and create a statistical model then we can start making predictions and then here i mean predictions future values like we can start making statements about what the expected value in in the future might be how the data or the behavior would look like and why do we want to do that so we can you know deal with uncertainty and uh that's uh and if you like from the uh at the beginning we talked about you know why um it's difficult sometimes to make decisions on data because there's some uncertainty involved and so we want to deal with uncertainty and we also once we can deal with uncertainty once we can you know make predictions say make statements about our data and future data we can also you know um say something about um you know normal good behaving data and abnormal not normal behaving data and why is that important because you know you want your business applications to run in a certain way you wanted you know your applications run smoothly without any headache and also in the future and if something happens you want to anticipate already before it happens and that's what servicenow also um does a lot you know they have predictive intelligence and anomaly detection systems uh and the reason why is because you know we want to anticipate issues and problems before they happen so yeah total sensor case sorry to interrupt but i think you you hit the nail on the header again because uncertainty um the the world of i.t operations is growing more and more complex every year so the way we work with applications the way we we work with business systems whatever it might be the complexity is growing and we see this a lot lately the past years with things like devops more and more monitoring tools you have cloud like all of a sudden you have all of these dependencies so when you have so many dependencies um then all of a sudden how i at least see it is that it becomes really really important to track what are the behaviors over time and like you mentioned seasonality um if we put that into a context if you are a company which which has like a web shop perhaps and then it's chris well we're almost in christmas now so a lot of people are buying things online then obviously you can expect a higher degree of traffic a higher degree of orders and then you have a seasonal trend so i think to become you know a next step in truly a data driven company then tracking the behaviors of applications is key so it's not only enough to you know react that hey we have a problem but you want like you explained intimately no how is the behavior of our applications how is the behavior of our infrastructure correct exactly you want to know that like you should know that you know your um you know cpu power for example or your computer servers are not working too much because there's like an um you know anomalous intrusion or there's an uh you know your computers are hacked and that's why it's using instead of like you know uh uh well you know that in christmas time it increases because it's christmas so but then in the right context that's why it's important if it was like in the summer uh you would know that there's something else going on like uh then you can look at it yeah right right and then you mention here before i cut you off triggering a normal alerts could you expand upon that because that's what it's all about right yes so uh at the end of the day so um when we make predictions um you know we want to business applications turn to run in a certain way and and when it comes to anomalies like anomaly detection systems so i i talked about that we train the models on training data and test data um but you know we also want to anticipate like you know future values out of sample like um so in this case when we talk about anomaly detection systems we are talking about the future anomalies that might occur and then what it's like it depends on the organization of the business so when the alerts are triggered either it's automatically like actions are automatically taken to solve the probable issue or the the alert goes to human operator would takes you know action or who looks at it and you know and then takes a decision based on these alerts and but so there's some coming back to the uncertainty and um i want before i give the floor to you alex um so there's some when we talk about the future events there's some you know mysteries involved but then we will talk that uh later on before we dive into um uh more on the anomaly section yeah so i i wouldn't like to really set this into a context here because there is there tends to be a misunderstanding when i speak to professionals especially in in the context of anomaly detection i always hear the term like if we compare the traditional monitoring the traditional way of working with events if we compare that to anomaly detection and i myself i was guilty of this in the past because i used to think that hey normal monitoring it it also warms when something is not normal so what is then really the difference between this more traditional monitoring ways of working and when we work with outlier detection and anomaly detection and this is one of the keys i think to understand that i'm gonna try to put it here in the context that traditional monitoring the thing there is that yes we can warn when something is not normal but the difference is that there are people who set those thresholds there are people who define when something is not normal and when something is normal so for example if the cpu is more than 95 for more than two minutes a person says that okay this is not normal so we are going to create a rule for it or a threshold in our monitoring system um but fundamentally it's still people who are setting these thresholds and there is like an endless tweaking of all of these different thresholds and the people needs to learn the systems manually over time how are the systems behaving how are the servers behaving and then you find yourself adapting these thresholds pretty much all the time and 10 years ago that was fine but like i touched upon previously when the sheer complexity has increased so much then manually adjusting these thresholds all the time um it's simply not realistic anymore and as an effect you miss a lot of different events that are happening within your ik infrastructure anomalies however um that's a bit different because then we are looking at really two things one is like we spoke about before the behaviors so how are the applications how are the metrics and how is the data behaving over time so that's when we use this time series database um and in a nutshell we are then creating a baseline of this is typically how you know the cpu is behaving for this particular server that belongs to this particular application for example and then when we see something which pops out which isn't really normal then we trigger an anomaly alert but the key here is that it can only indicate whether something is abnormal or normal it cannot really say if it's good or bad we still need a person to indicate that um and i don't think that's gonna change for a while um and then we have the predictions so this is what what you spoke about keith when we really are predicting future events based on the historic data and we can create these sort of predictive alerts and i'm gonna try to explain this so um like the le5 um explained to me like i'm five so what's what's the deal with predictions really well if we're looking at this graph here it could be like a graph on the cpu or it could be of transactions or whatever it is we see the blue line and that's the things which is currently happening and then we have the orange line which is what do we predict will happen and because we have time series data because we have all of this data from the past we can make fairly accurate predictions or hopefully at least um of yeah how will the coming five minutes look how will the coming 10 or 30 minutes look because we have a lot of history so we know that last time this sort of event happened last time um whatever happened then you know the the iit infrastructure behaved in a certain way so here i'm going to predict that behavior again so essentially we are detecting anomalies in the future which have not yet happened potentially so i just wanted to kind of debunk these concepts and get it out and declare there so what is the difference between traditional monitoring and anomaly detection and why is it so important so this is truly what it means to be working according to predictive models um but you you have more to say about the predictions akif so yes there is a mystery around it because it's not just that simple correct there's a mystery on the predictions because they're i mean like you know um you hope to capture actually in your model all the relevant data all the factors that are relevant to make good predictions right so you want to know all everything actually on one hand because you know like you said you know when you make the prediction whether it's like gonna be uh you know the anomaly or not and uh you showed in that nice um uh i mean what if you have like in historical data where you haven't seen any um anomaly like you have to learn new models you know what an uh anomaly or how enormous behavior you know should look like um but you know the world is complex and um and there are lots of you know um that's why there's some some mystery involved in predictions especially like you know uh future uh you know predicting the future um but first of all like why are predictions important i mean you want to you don't want to be in reactive agent you don't want to be reactive but like more proactive like anticipate the future and you want to you know predict anomalies and why like you want to handle it before you know it becomes a problem and um this is important to stress and uh otherwise then uh yeah it doesn't make sense you know uh all the tools and you know nice uh tools you have and and then you want to know which or normally will really cause the problem because um i mean when we talk about like outliers and uh or normally like you can have an outlier because you know data is not like constant it has like it's varying like get lots of different kind of data you know and then you try to fit them you know a model like you know nice machine learning statistical model um mostly like average in the value but i mean there will be you know data points um that should not be you know be in there when you construct your model so when you create your model um there will be anomalous behavior or anomalous data in there and you want to detect that and learn about it and then so you can prevent it in the in in the future so what you're saying is that the success of these statistical models depend on you capturing when things go wrong when when things are not as as we expect them to be and if you do that over time you kind of learn that behavior of this is how it looks when it goes wrong this is how it looks when when it's normal and but but the success is really depending on when a crash happens then you need to capture that data exactly exactly and we will talk about you know uh that's more uh but this is more in an ideal you know situation so when you have ideal statistical model it will do that it will look at you know your data it will say you know it's going to behave like this you will have no problems at all and you can you know run your it operations very smoothly uh without any head heck so what happens then keith because i think that's the next slide what happens then when we have problems that never have occurred before that are completely new to the statistical models and everything well then we get uh you know a black swan event we call it uh you know it was a coined by nassim taleb and then actually this is also the definition of mystery so if you haven't you know captured the factors or like the data you know uh for such an you know event uh i mean it's that's why uh we called it like the mystery of predictions not a puzzle of predictions like if it was a puzzle we know there's some solution answered there we just have to go and you know look for it and find it but you know when you when we talk about the future there are like factors and data that are like not known i mean there are no unknowns and and there are unknown unknowns and especially the unknown unknowns uh like uh great guy once said it uh you know the biggest catastrophic events come from there and i also want to say in uh you know uh for example when you create like statistical models and machine learning models it has a certain accuracy probably 95 percent of the time they will be you know right and or even more like we can create really high accurate models but still um you have to capture you know that data that factors that could cause you know anomaly um and and then you know take action that's the second part you know take action when you get the anomaly alert for hey there's going to be a problem um so coming back to that mystery part so um yeah like i i the moral of this is story with the medieval knight i wanted to show actually like uh so you know you have nice armor you are prepared uh but there's you know small like opening to see uh and then you know you get a good uh bowman will you know uh hits like directly you in the uh you know in that small opening that's there yeah so my question always is like you know or the moral of the story is what is the probability that the error goes through that opening i mean one percent maybe two percent five percent or maybe zero point one percent um then i then always you know state that murphy's law that if anything can go wrong will go wrong very pessimistic view um especially if it's a you know systems created by humans uh i mean we also create solutions for it so yeah so a question here because i'm i'm sure that some of the audience is thinking the same what i am thinking um okay it's you know if we look at how problems occur in the it infrastructure if we look at how issues occur in applications or in our business services or whatever it is if we have these issues and we capture them over long enough time then we can in a fairly accurate way predict that aha we have seen this pattern before and now we we start seeing the same things repeating again so like you mentioned with 95 accuracy yes this is going to happen within the next 10 or 15 minutes for example and that's great but often the just like you said the unknown unknowns um are the things which um creates the biggest problems in it infrastructure that creates the biggest problems when we speak about running a business it's the thing that you really that have never happened before and that is very difficult to anticipate so in your point of view what would the future solution to this be and are we there today can we cover those black swan events um well it's a you know let me come back to the definition of black swan event it's that rare high impact like um you know prospectively you know in prospective it's not predictable like um that's what i was trying to actually say they will always you know um uh once in a while you know um a problem in issue uh you know that will have a high impact but you could have you know said okay this could happen like let's say your service with all of them at once you know shut down yeah it could but then the chance is like 0.001 and maybe like i don't know if it's a good example like so we had that with you know a couple of weeks or months ago with facebook and whatsapp and uh yeah i asked myself how could this happen like i couldn't you know my whatsapp and facebook and you know it didn't work so um but it can happen and that's not like happening often or hopefully not happening that much many times and which is difficult to capture and the moral of the story is you should have a strategy rules and procedures in your decision making in place if in search in a case it happens uh i mean you should we can trust the ml and ai algorithms like uh to a certain extent but still like um you should have like a strategy in place in case of such an event happens you know what do you have to do like um how should you you know react um and then yeah i mean with predictions uh i mystery like you know that that's why it's mystery of predictions and not puzzle of uh you know yeah it's a puzzle and um and then miss you you have to you know uh deal with it like so what you're what you're saying is there needs to be a process around it that if we can anticipate we can forecast we can learn you know the best behaviors of our systems and to maybe a 95 to maybe a 99 accuracy if we're really good but it is that one percent which we will probably never be able to capture that you also need some fallback methods for exactly yeah but i think the the general you know the consensus here is that predictions are really really good because then we are working based on historic things we we are working based on facts on data but they are at the same time not a silver bullet in all the cases and i think those are fair expectations to set when we speak about aiops and predictions that not everything can be predicted for the reasons you guys mentioned exactly and another thing just to um yeah finish it like it's um so business obligations like the data from it it's not like from you know uh the stock market which is much more you know complex lots of other factors i mean tests you know behave in a normal way most of the time so yeah so it's a good point actually it's a good point most of the time data is behaving normally um with that being said we only have 10 minutes left and i we have now discussed the the theoretical concepts quite a lot so we have been speaking about the pyramid of intelligence what is the difference between machine learning and ai what is time series and how do we capture it and how do we apply statistical models to it and how do we create these baselines and how can we try to forecast the future so this all sounds good on paper but a lot of questions a lot of people they they often tend to ask the question like okay that's that sounds good but how how can we get there um so that's what i wanted to stress here is that when we speak about aiops when we especially serve is now item and machine learning um it's of critical importance to have some tangible use cases um so this is something that we have worked with with other organizations about and but as a first step so again this this is just to give the guys and girls on the call and an understanding of how could we do this ourselves how could we start this journey in a very simple way so first of all the sources where do we get our metrics from that we will store in a time series database um like you mentioned actually we can use the agent client collector of service now which kind of scraps the matrix um the metrics and the matrix um and then we also have things like traditional application performance management tools that we can tap into metrics from but then it's important because when we start with this i mean there could be potentially hundreds of different metrics that we could measure and create statistical models for but that's not really valuable but rather we should select a few critical metrics that we think are important and maybe also for critical business applications so minimize the scope a little bit and focus on really learning the behavior and for a few metrics and maybe 10 20 30 rather than hundreds of them um and then do that for a set of yeah very important applications because then we get realistic we learn this intimate behavior like you explained before that [Music] but then when we speak about use cases so how can we pitch this how can we how can we refer to these things to management or if we want to get like a buy-in from a vp or whatever it might be we need to have tangible use cases so yeah just knowing the statistical models of our metrics that's cool and all but what are we gonna do with it so that is where we come to that act part um in the pyramid so i would say four very common use cases is first of all moving away from static thresholds using ai ops and using these sort of machine learning is a great way not to having to rely on static thresholds because that brings me to the next point we become aware of the behaviors of our applications and when we are aware on those behaviors we can start analyzing anomalies and more importantly we can start creating anomaly alerts and we can act on these anomaly alerts sometimes it's fine um like nothing bad has happened but other times these anomaly alerts can be super um good important critical business data to be aware of that like hey we have anomalies in our system we need someone to look into it before a user calls or before a monitoring tool warns about the cpu um and last but not least to connect these anomaly alerts to the normal itsm processes so creating incidents out of them um maybe working with change management etc it's also very realistic things um so of course we have only scratched the surface here now a little bit um but i i wanted to end the note here with that it's not just a theoretical world what we're speaking about here but this is available commercially off the shelf realistic things to to get started with even if it's in a small scale proof of concepts it can be done and i think that is is super cool so let's see we've we've received some questions um i see you're also typing an answer there keith um but then i'm gonna i'm gonna read out some questions now here so the first question is actually from from stayner and he is asking how far are we going to go and predict what might happen and what is the cost of predicting um so i'm gonna quickly try to answer the first part of that question how far do the predictions stretch and i think that depends on the model if i'm not mistaken akif and we can also influence this in service now but most of the times it's like 30 minutes and and the further you go the more inaccurate the prediction becomes correct yes exactly and depending on your data too so if you have for example two weeks of data i mean uh then you can look at cert you know set up a certain like window uh to uh you know how much you further can look into like right let's say if it's two weeks then maybe like one to two weeks into kind of the future uh depending on the model and the data actually and that that's one of the factors you have to actually consider uh in your strategy especially when you have um and you want to install like um predictive intelligence and predictions yeah and that's one of the critical factors so so basically the more data you have the further you can predict in the future to like summarize it um and then there was also a question about the cost element so what is the cost of predicting this um i i don't have the exact like if you're asking about the license cost etc um i don't know that but generally speaking the cost of the predictions are fairly low um like again you got to put this in a context but let's say implementing this in a small proof of concept it's a fairly quick project it doesn't take too much time in my experience um so the cost of the predictions the cost of storing all of this data is relatively low compared to only like five ten years ago so that's what makes it a rather interesting part um there's also another question which is how the variables of a time series are identified and how can they be corroborated yeah um so that's a very good question too um so um i mean like when you get cpu data and all the data you know um so you can do it in two ways um like you know you can beforehand theoretically set up like the variables you know factors you say this influence this uh you know um you know the data how the data is generated uh you can do it in a non-parametric way you know let the data decide actually these uh variables or factors or parameters in the statistical model so yeah build two approaches in identifying these variables excellent thanks again and then there's a final question from ika which is if we there's any implementations with dynatrace so far so even though we're speaking about servicenow item here today it's very common that we um hook into data from dynatrace because dynatrace is um is monitoring application and performance data so that's that's a very common source actually so with that being said um i said we have one other question from katarzina um you have given a proper reply there i don't have time to read it out loud but i'd like to thank the audience so much for listening to us here today indeed it's a rather technical master class but um that's the future it is a future of data and technology so i hope that we we we did a good job in trying to explain the high level concepts but if you on the call are interested in knowing more about these sort of topics um you achieve have written a very extensive report that we're going to publish tomorrow more about predictions um so everyone on the call here you will receive a link to achieve report but more importantly you can always reach out to us so it's just ping at another partners if you have more questions around this even if it's just like a technicality you want to discuss it doesn't need to be a project and of course stay in the loop so follow us on linkedin follow us on our youtube channel and we will make sure to post more and more content in this area and i'm certain that we will see your face again i keep so thanks a lot for joining me here today yeah so thanks uh thank you alex uh thank you and excellent really great insights into uh to this from the business perspective perfect then thank you everyone and have a nice day and we'll be in touch

View original source

https://www.youtube.com/watch?v=wgP5rC6TroI