Predictive AIOps - Technical Deep Dive
hi my name is ben randolph and i'm an advisory solution consultant as organizations make a push toward automation to modernize the management of it we'd like to discuss and demonstrate how servicenow's predictive ai op solution can help you achieve your goals as a technology leader you and your team have to balance a nimble environment while doing more with less predictive ai ops helps facilitate that and as a matter of fact like our cio's initiatives at servicenow many cios and leaders like yourself have similar imperatives around the pursuit of automation within the infrastructure and how to reduce outages to deliver always on services so in the vein of doing more with less driving your footprint to the cloud so zero reduced physical footprint to reduce cost and transfer risk is important and then effectively designing your its state so there are zero outages which translates to reduction in loss of revenue or service degradation and reducing financial impact and then ultimately enriching and scaling your its state through frictionless stable work using ai machine learning and letting the machines do the work for a team like yours to automate away incidents you'll need to leverage itsm pro to tackle people-generated incidents and itom enterprise to tackle machine-generated incidents especially as it relates to predictive ai ops for the sake of this presentation so did you know that for an average enterprise like yourself approximately 60 percent of incidents are generated by machines and then forty percent of incidents are generated by people the machine generated incidents are coming from various solutions that are managing your company's operations from the i t side and from the infrastructure side so together the people generated and machine generated incidents constitute all the level 0 level 1 level 2 and 3 incidents that you would face so at a macro level employees and users and end users are demanding self-service and much faster turnarounds and it wants to reduce tickets and resolve issues more quickly so the machine generated incidents are exploding with the increase in complexity and proliferation of the it and technology stack as we look at the ai ops journey maturity model or maturity curve most peer organizations tend to fall into levels we'll say one and two there are still manual processes at this level wall rooms to determine root cause point solutions for monitoring and so on and having said that virtually every it organization wants to progress their maturity further with additional insights and proactive automation capabilities progressing into levels three and four requires companies start to take advantage of a single data platform and cmdb to unite it operations data they can reduce noise and false positives allowing them to do automated root cause analysis they also can correlate it operations data to create recommendations in solving issues which reduce mean time to resolution furthermore they have a feedback loop with devops teams or infrastructure teams knock teams itsm teams for high risk changes for example the organizations in levels three and four are using ai ops with itsm and operations management to glean insight on predictive problem analysis and eventually change risk analysis at this point organizations are using automation playbooks using them to automatically resolve common issues and then ultimately eliminating outages and incidents altogether where would you say your organization is on this maturity curve so as organizations strive to leverage a more predictive and ai driven operational model there are three primary tenants we'll discuss throughout this presentation first the predictive ai op solution predicts issues by analyzing log data and before problems occur immediately the platform uses ai machine learning to determine root cause and then quickly give you early warning into service impacting issues before the business is impacted second we'll correlate the context from log data and incoming events and related to a service aware operation to prevent service degradation and then enact remediation and self-healing by integrating with the tools within the servicenow platform like flow designer app engine and then ultimately tied that back itsm and change management processes to streamline the flow and process overall and ultimately that value that real life the realized outcome is a prevention of 35 percent of p1 outages improvement of mean time to resolution by 40 so as you think about that we think about inherently log data is designed for troubleshooting right so typically 95 percent of issues can be identified by searching through log data what we are positioning here is a way of changing our customers traditional monitoring workflow and what we often see today are i t teams working with monitoring solutions setting thresholds and trying to write the actual logic to foresee any potential alerts and a no and notify the team and then often defer to the logs anyway to further troubleshoot but this entire approach is reactive by nature so we propose a paradigm shift rather than having to tell a machine what to look for through setting manual thresholds and things like that the machine will learn the log patterns from your different sources and tell you where issues of interest exist so as we show organizations how they can leverage the power of the servicenow predictive ai ops solution we want to help you and your team understand how they can transform the i t experience and gain a competitive edge using business context and machine learning to avert outages and resolve resolve i t issues proactively so let's take a look at our standard grouping of it business services and you can probably relate to this there is a mix between essential trivial and of course critical applications that drive revenue and customer consumer facing services then there are event sources enrichment data that come from various sources like your cloud infrastructure on-prem servers storage etc these event sources produce data that are then forwarded to our event management solution and this event driven data can come from a variety of sources anything from familiar monitoring tools like solarwinds or native cloud monitoring tools as well as log aggregators with all of this rich infrastructure data this is where our aiops engine really shines the platform is able to help your team achieve 90 noise reduction because disparate events about cis are correlated to actionable alerts and root cause is automatically highlighted along with impact analysis this is helpful to know how a business service is being impacted and how any changes may affect other services in your environment using statistical models and the existing event data metric anomaly detection facilitates the ability to predict that a ci is at risk of causing an outage furthermore the advanced correlation function allows incoming events to be correlated to a grouped alert and prevent swivel chair problem resolution that many it teams face with our solution you have actionable information in one workspace view and collectively with our extremely powerful log analytics engine our solution is far more powerful and comprehensive and differentiated from our competitors one because alongside the other elements in the platform we are able to ingest log data from disparate sources to preemptively and proactively draw conclusions of imminent problems within the environment at which point operations teams can take action before an issue occurs in your infrastructure and there is real financial time-saving value here so the predictive ai op solution gets you out of a reactive state by predicting problems before occurrence and quickly preventing outages by identifying root cause and achieving significantly lower p1 p2 incidence and ultimately that leads to a 40 reduction in mean time to resolution and as we discussed earlier progressing along that maturity curve towards automation and preemptive problem resolution by the system so that journey towards a self-healing data center let's dig into the predict prevent and automate functionality a bit more by starting with predicting issues before they occur in your environment so how does servicenow's predictive aiop solution work by extending and augmenting log management every system produces logs so we can leverage that data to help our customers predict issues first the system starts with the collection of raw log data in real time and manipulates the data feeds via auto parsing this can be done via passive input like our syslog or our htc or agent client collector or agentless native data inputs from amazon s3 or kafka for example or even log forwarding from raw data sources after collection the predictive ai app solution will start to model the data and understand normal behavior the time horizon will become broader after implementation so one hour then one day one week etc and normality will become understood over those time horizons the important element in this step is that the predictive aiops engine will advise on deviations from learned patterns and then the event management correlation engine will correlate the anomalies from the log data with other existing events in the platform and associated interrelated cis and services anomalies in the logs could be recognized as something occurring above or below an expected pattern like a certain average number of order completions during a given time window or an anti-pattern was recognized instead of a spiking activity which could trigger a threshold alarm or traditional threshold alarm a significant drop in logins or processing utilization would be an example of anti-pattern anomaly and that also will be captured and associated within the solution additionally the platform will recognize a net new pattern as a potential anomaly the takeaway here is that this detection is all unsupervised and the platform can recognize normal language in logs and text patterns combining all of this intelligence we are facilitating root cause in your environment second the tenet of prevention is an avoidance of service degradation and or loss of revenue or functionality within your organization so the platform leverages the incoming event driven data along with metric data to garner insight for preventive problem resolution and leveraging existing knowledge based articles stored within the platform which include thousands of pre-loaded and customer-fed insights as well as ai drawn root cause analysis technology organizations like yourselves are equipped with the ability to drastically reduce negative impact to the infrastructure and eliminate outages finally once a workflow for remediation is created the platform can take action open an incident proactively and then referring back to that maturity curve get to the place where your organization is poised to transition to a self-healing state based on those triggered actions resolves issues in the environment via automated workflows and automatic tasks and as your organization becomes more comfortable with automated problem remediation the functionality within the solution is there to facilitate so finally when we discuss automation especially automation across team workflows when forging towards automation too often it operations teams have to focus on their individual silo when needing to resolve an issue and they may have to ultimately interact with a service management team or a security team or development team and use phone calls and emails and instant messages and things like that to kick off manual processes so let's think about a typical workflow in your organization if your operations team identifies a problem and completes root cause analysis typically you might need to create a ticket and figure out who should work on it complete the remediation and then close the feedback loop when you traditionally deal with multiple teams and tools that's when new problems and added friction can occur and with the power of the of the now platform you can automatically create and automatically route the incident to the right group with intelligent categorization as well as assignment you can further simplify tasks by leveraging pre-built playbooks and using no-code workflows to further automate the process so when the actions are taken the system has all the pertinent information to resolve an issue preemptively and then close the feedback loop as your organization is more comfortable with the predictive ai app solution and automation the possibilities become endless by leveraging our low code and no code tools like automation engine app engine flow designer and the like ultimately the servicenow predictive ai app solution and now platform facilitates business collaboration and problem avoidance across various teams and stakeholders inside and outside of your organization the ability to detect where there are changes or problematic alerts and alarms coming from existing solutions and then immediately distill and prioritize next steps is key the servicenow platform is the foundation for your organization's digital transformation as you are able to quickly identify and automatically assign tasks to the right team and as we discussed pinpoint root cause reduce mttr and proactively resolve issues through the use of predictive ai ops all driving improvement to your technology operations and user experience so here's a real world example and some customer feedback with respect to the servicenow predictive ai ops solution like most retailers one of our customers a nationwide retailer had to move their brick and mortar business online overnight because of the pandemic and their traditional ai op solution did not pick up on a payment application slowing down until it reached the threshold they had manually established for it servicenow's predictive ai ops solution was running in parallel at this customer at the same time and it picked up the slow forming anomaly 78 percent faster than the customer's traditional ai app solution saving that customer 1 000 orders and reducing the time to repair the issue by 40 so in this example unlike traditional noc or ai apps tools which only deal with known failure scenarios predictive aiops uses advanced machine learning to reveal what customers don't know uncovering complex unforeseen issues that can lead to future services outages and impairments the predictive ai app solution leverages the power of machine learning artificial intelligence and automation to predict service issues pinpoint the root cause and automate remediation much faster it replaces the tidal wave of monitoring data that you all are familiar with with a small number of actionable alerts helping to prevent service issues and resolve them rapidly so let's take a look at how smooth the deployment of predictive ai ops is and then we'll dive into a brief demo to show how we can achieve what we've proposed as you consider what we've discussed in the power of the now platform and the predictive ai op solution it's important to understand that the solution was designed for ease of implementation and quick time to value so we're talking weeks not months or years here's a quick snapshot of how that's possible first your team would prepare by identifying a few of your most critical services that would be in scope and would achieve value by preventing any future outages to that service in week one you would connect predictive apps to existing log aggregators for example splunk or sumo logic or elasticsearch in other words we tap into your existing log repositories and then you're off and running it's easy in weeks two and three we will optimize out of the box ai machine learning algorithms to read and understand the log data which tunes to your specific environment and by week four or sooner your teams are ready to deploy to production and achieve value and expand to onboard more services to monitor also note that the predictive ai app solution can be leveraged at any state in the cmdb health or maturity as long as the cmdb has application service registered for the application any other service that you want to monitor predictive ai ops will work and will enrich the cmdb by automatically identifying cis from logs so we see this as a win-win for you predictive a ops has a fast time to value and it enriches the cmdb and that's it if there are further questions and interest please engage with your servicenow representative so we can work with your team and plan for a deeper discussion and predictive ai ops workshop thank you
https://www.youtube.com/watch?v=iavTyNQtz1o