AI-Powered Service Operations Part 1 - deflect, predict and notify
hi this is evans nicholson from technical marketing at servicenow today i'm going to talk to you about how we can deflect and predict issues before they affect employees it teams and customers outages and service degradations are stressful times for it teams including knock operators agents as well as employees as we've all learned these outages though preventable are inevitable given the complexity of modern applications ai powered service operations enables the move away from an incident creation culture to more of an alert based actionable mindset this helps you find problems and solve them in a more predictive fashion versus being stuck in problem change and incident management the key is being proactive versus reactive and being collaborative and not siloed in this area servicenow can help in four ways first we can help identify if all the events coming in are tied to a single problem with event correlation a second we can find the root cause of the issue faster third when we see this we can actually notify employees of the disruption and the fact that it's being worked on so they don't keep opening tickets and fourth we provide remediation suggestions and actions to fix the problem in this demo we're going to focus on the first two outcomes i'll cover the other two in another video first let's start with amelia bryant our service desk operator amelia is seeing a flood of new tickets entering her queue from one of the tickets we can see servicenow predictive intelligence at work identifying 19 similar high-impact incidents and agent assist is recommending we open a major incident based on this grouping and if we drill down further we can get a great view of the impacted service and we can see we have one critical incident open and looking at the timeline we see more activity recently than in the past which is a further indication of an ongoing issue there are 89 tickets open for this service alone so obviously something bigger is going on so this is a quick look at the itsm workflow but let's take a look at itom and see what our itom team is doing in parallel now i'm going to switch over and look at this as roberto hopper who is on the operations team at roberto and amelia's organization these are separate teams handling the two responsibilities but this could easily be the same person depending on how things are organized roberto starts on the operator workspace and sees that the order status application is in a critical state we can see that itom event management has combined a group of related alerts making it easier to troubleshoot as one event we can see a number of alerts coming in from many sources and notice that this list includes alerts from other vendor solutions this is one example of what we mean when we say that servicenow is the platform of platforms the first alert that came through is from health log analytics hla which is part of itom health continuously monitors logs to identify any anomalous behavior this is predictive ai ops at work i say predictive because you can see this hla alert came in several minutes before alerts from the other sources changes in logs are great preemptive indicators of issues to come in essence what we're seeing is a timeline of events around a service degradation or failure and it all starts with hla so let's take a look at this hla alert the onscreen mouse over gives us context into what we're looking at there is a sudden anomalous increase in the number of logs as compared to the baseline remember what's important is not only what's in the logs things like failure resource not found fault things like that but also the amount and the size of the logs and that's what we're seeing here now we'll go back to the group alert and see if we can help roberto determine root cause analysis in the details pane we'll scroll down to find the primary alert in this group it looks like we have a disk space issue on a windows server and you can see here the related cis and criticality in the essence of time i'm going to skip right to the root cause analysis conclusion but if i wanted to roberto and i could look at the metrics section to get an in-depth look at stats and performance on this windows server as well as the associated applications or services but luckily root cause analysis has done a lot of that math for us and here we see two probable root causes of the service issue the bottom one appears to be a planned devops change and the top one is an unauthorized change with the information he's seen so far roberto decides to officially trigger a service degradation this will inform the employees who rely on the order status process and the porsche online service that there's an ongoing issue we've just seen how servicenow can help it teams solve complex issues faster and more collaboratively we can help identify if all events coming in are tied to a single problem with event correlation and we can find the root cause of the issue faster in part two of this video i'll cover how we notify employees of ongoing issues and how we provide remediation actions and suggestions to fix the problem for more information check out the main itom product page thanks a lot and i hope to talk to you soon you
https://www.youtube.com/watch?v=Zmbevk-YfNU