Predictive AIOps - Predict, Prevent, and Automate
so as organizations strive to leverage a more predictive and ai driven operational model there are three primary tenants we'll discuss throughout this presentation first the predictive ai op solution predicts issues by analyzing log data and before problems occur immediately the platform uses ai machine learning to determine root cause and then quickly give you early warning into service impacting issues before the business is impacted second we'll correlate the context from log data and incoming events and related to a service aware operation to prevent service degradation and then enact remediation and self-healing by integrating with the tools within the servicenow platform like flow designer app engine and then ultimately tie that back itsm and change management processes to streamline the flow and process overall and ultimately that value that real life the realized outcome is a prevention of 35 percent of p1 outages improvement of mean time to resolution by 40 so as you think about that we think about inherently log data is designed for troubleshooting right so typically 95 of issues can be identified by searching through log data what we are positioning here is a way of changing our customers traditional monitoring workflow and what we often see today are i teach teams working with monitoring solutions setting thresholds and trying to write the actual logic to foresee any potential alerts and a no and notify the team and then often defer to the logs anyway to further troubleshoot but this entire approach is reactive by nature so we propose a paradigm shift rather than having to tell a machine what to look for through setting manual thresholds and things like that the machine will learn the log patterns from your different sources and tell you where issues of interest exist so as we show organizations how they can leverage the power of the servicenow predictive ai ops solution we want to help you and your team understand how they can transform the it experience and gain a competitive edge using business context and machine learning to avert outages and resolve resolve i.t issues proactively so let's take a look at our standard grouping of it business services and you can probably relate to this there is a mix between essential trivial and of course critical applications that drive revenue and customer consumer facing services then there are event sources enrichment data that come from various sources like your cloud infrastructure on-prem servers storage etc these event sources produce data that are then forwarded to our event management solution and this event-driven data can come from a variety of sources anything from familiar monitoring tools like solarwinds or native cloud monitoring tools as well as log aggregators with all of this rich infrastructure data this is where our aiops engine really shines the platform is able to help your team achieve 90 noise reduction because disparate events about ci's are correlated to actionable alerts and root cause is automatically highlighted along with impact analysis this is helpful to know how a business service is being impacted and how any changes may affect other services in your environment using statistical models and the existing event data metric anomaly detection facilitates the ability to predict that a ci is at risk of causing an outage furthermore the advanced correlation function allows incoming events to be correlated to a grouped alert and prevent swivel chair problem resolution that many it teams face with our solution you have actionable information in one workspace view and collectively with our extremely powerful log analytics engine our solution is far more powerful and comprehensive and differentiated from our competitors one because alongside the other elements in the platform we are able to ingest log data from disparate sources to preemptively and proactively draw conclusions of eminent problems within the environment at which point operations teams can take action before an issue occurs in your infrastructure and there is real financial time saving value here so the predictive ai op solution gets you out of a reactive state by predicting problems before occurrence and quickly preventing outages by identifying root cause and achieving significantly lower p1 p2 incidence and ultimately that leads to a 40 reduction in mean time to resolution and as we discussed earlier progressing along that maturity curve towards automation and preemptive problem resolution by the system so that journey towards a self-healing data center let's dig into the predict prevent and automate functionality a bit more by starting with predicting issues before they occur in your environment so how does servicenow's predictive aiop solution work by extending and augmenting log management every system produces logs so we can leverage that data to help our customers predict issues first the system starts with the collection of raw log data in real time and manipulates the data feeds via auto parsing this can be done via passive input like our syslog or our htc or agent client collector or agentless native data inputs from amazon s3 or kafka for example or even log forwarding from raw data sources after collection the predictive ai app solution will start to model the data and understand normal behavior the time horizon will become broader after implementation so one hour then one day one week etc and normality will become understood over those time horizons the important element in this step is that the predictive aiops engine will advise on deviations from learn patterns and then the event management correlation engine will correlate the anomalies from the log data with other existing events in the platform and associated interrelated cis and services anomalies in the logs could be recognized as something occurring above or below an expected pattern like a certain average number of order completions during a given time window or an anti-pattern was recognized and instead of a spike in activity which could trigger a threshold alarm or traditional threshold alarm a significant drop in logins or processing utilization would be an example of an anti-pattern anomaly and that also will be captured and associated within the solution additionally the platform will recognize a net new pattern as a potential anomaly the takeaway here is that this detection is all unsupervised and the platform can recognize normal language in logs and text patterns combining all of this intelligence we are facilitating root cause in your environment second the tenet of prevention is an avoidance of service degradation and or loss of revenue or functionality within your organization so the platform leverages the incoming event driven data along with metric data to garner insight for preventive problem resolution and leveraging existing knowledge based articles stored within the platform which include thousands of pre-loaded and customer-fed insights as well as aiadron root cause analysis technology organizations like yourselves are equipped with the ability to drastically reduce negative impact to the infrastructure and eliminate outages finally once a workflow for remediation is created the platform can take action open an incident proactively and then referring back to that maturity curve get to the place where your organization is poised to transition to a self-healing state based on those triggered actions resolves issues in the environment via automated workflows and automatic tasks and as your organization becomes more comfortable with automated problem remediation the functionality within the solution is there to facilitate so finally when we discuss automation especially automation across team workflows when forging towards automation too often it operations teams have to focus on their individual silo when needing to resolve an issue and they may have to ultimately interact with a service management team or a security team or development team and use phone calls and emails and instant messages and things like that to kick off manual processes so let's think about a typical workflow in your organization if your operations team identifies a problem and completes root cause analysis typically you might need to create a ticket and figure out who should work on it complete the remediation and then close the feedback loop when you traditionally deal with multiple teams and tools that's when new problems and added friction can occur and with the power of the of the now platform you can automatically create and automatically route the incident to the right group with intelligent categorization as well as assignment you can further simplify tasks by leveraging pre-built playbooks and using no-code workflows to further automate the process so when the actions are taken the system has all the pertinent information to resolve an issue preemptively and then close the feedback loop as your organization is more comfortable with the predictive ai app solution and automation the possibilities become endless by leveraging our low code and no code tools like automation engine app engine flow designer and the like
https://www.youtube.com/watch?v=p6xYPwu8580