logo

NJP

Demo - Quebec - Predictive AIOps with Health Log Analytics

Import · Mar 03, 2021 · video

hello i'm mark rozinski a product success architect at servicenow and today i'll showcase how servicenow's predictive ai ops with the use of health log analytics can help your enterprise predict and prevent i.t issues in real time we start in the operator workspace here we see that one of our most critical application services e-banking is experiencing a potential critical issue let's take a look on selecting the application service i can see that there are currently eight critical alerts let's click on the service map to get a better understanding of the situation in the service map we see a visual representation of the infrastructure that comprises the e-banking application service i see that the h a proxy as well as the two tomcat servers are impacted in red the proxy appears to have the group of alerts in the situation and is the first place that data enters the application service so let's start there in the group alerts screen i see the primary alert error and configuration file with a very specific configuration file mentioned as well as the additional alerts but to view all of them let's click the view more button here we see log analytics and our traditional alerting tool zabx infrastructure monitoring xavix is reporting on traditional metric based anomalies as we can see here and log analytics is actually in advance of xavix by up to over 10 minutes where we're reporting those specific configuration file anomalies five minutes after the configuration file errors we see that user transactions can now be completed and that the connection pool is full so we have greater context from these log-based alerts and only later does zabic start to alert us so here we see the proactive nature of the log-based anomalies and the added context that could potentially help us drive faster meantime resolution let's click into this log analytics alert which is the very first one reported to see if we get even more information here we see the exact log line errors found in the configuration file and on the right hand side we see how the anomaly behaves so typically this was inactive or not seen at all and then we clearly see a step up here where we're regularly seeing these exact types of errors where this configuration file is generating errors and quite significantly below that we have additional key value pairs extracted from the logs to give us greater context as to what is happening in this particular case there are different types of connection counts but there's only one host linux 301 if we look at that host it's a configuration item and that configuration item combines with impacted application services and there are two in this case not only e-banking but also linux servers as this configuration item is part of both application services if we want to investigate about the surrounding logs we certainly can as many troubleshooting experts would so let's view the surrounding logs and here instead of having to ssh into a server look for that particular file open the file and then look for the particular time frame where the issue occurred or is occurring we have that all done for us here we have the raw log line on the right hand side and a more visually appearing representation of the message here in the middle if we want to start slicing and dicing these logs further maybe searching for additional key value pairs or filtering for specific hosts we certainly can do that in our dedicated log viewer within the platform so we have added context around this anomaly but what can we do to resolve it for that let's go back to the overview screen and open up agent assist if we search our knowledge base which is done automatically we can see that someone has already run into load balancer issues before and if we click on the article we have a friendly configuration file directory where we can go to see what type of resolutions we can make to this particular configuration file okay so we know that we can take action on it as well and hopefully take action on it before any customer tickets are made so now let's look for a potential root cause so if we go back to the group of alerts we see here that there is a probable root cause tab from the probable root cause tab again we are using the configuration item we found in the logs to correlate to a change request selecting that change request we can see that this is related to a tomcat configuration on the load balancer if we scroll down lower we see that this is just minutes before the initial configuration alerts generated by health log analytics began we can be fairly confident this plan change is causing the alerts within the e-banking application at this time so now we have a root cause as well if we want to take it a step further and turn that knowledge base that we had into a automated remediation we certainly could using flow designer and it would be run here from the actions tab with this we've covered the journey of predictive a ops and responding to an alert before customers are affected if you have any questions or further interest make sure to reach out to us here at servicenow and we would be happy to help thank you for your time

View original source

https://www.youtube.com/watch?v=WQYfWHtJLrQ