ServiceNow AIOps Platform - Onboard New Apps
hello my name is Jason Smith and I'm the outbound product manager for item health and servicenow today we're going to go through how to use machine learning to onboard and manage applications with the apps platform okay so what we're looking at here is a very simple dashboard and the purpose of making this dashboard was to help us manage the crowdstrike installation of the agent across different machines throughout the infrastructure so in this case I can see that there is one active ola there's four minutes left and then one minute that has elapsed since the uh since the oil light got kicked off this is essentially telling me that I've got four minutes before an incident is going to be created we have some metrics here so in this case this is for the CPU this is really impact from the crowdstrike agent itself on the local machines and also memory and then there is some log data that we're able to get and pull into this dashboard also in addition we have these uh alerts that we can take a look at so the overall service has gone critical there's one open alert but how did we create this dashboard and what do we have to do to onboard the application so first of all um there was no pattern or a way for us to identify the Falcon sensor that agent from crowdstrike on its own there was an out-of-box content for that it's an agent but thankfully we're using Discovery and it has a feature called application fingerprinting so it was able to identify the application for me and what I did was I clicked on this found Falcon sensor and then ultimately just pressed submit this is one of the results you can see here so it gave things like a suggested group name I could view the processes that were ultimately being used this is all correct this is on Linux it's running Falcon d and then if we go back to the this page here we can see that uh when I had I mean is essentially pressed submit before it created a discovery pattern for me and also a CND class very importantly it created something called a process classification let's take a look at the process classification so this was created through the application fingerprinting and it says that um basically when Falcon D is detected on a machine it's going to kick off this probe which will go and identify the the application running on the machine it's very worthy to note that this process classification is also used by the agent client collector so when agent client collector detects this process running on a machine that application will automatically be entered into the cmdb in addition I will need some policies so I'm looking now at the agent client collector policies and what I'm trying to do is bring this application under monitoring so I created a policy and this policy is monitoring see eyes of the type crowdstrike happens to be running on three different agents and then I have these check instances so let's take a look at those this first one is just to see if the process is running or not and here's the configuration all I had to do was say it it would be critical if there is zero processes named Falcon ID running so if it's under one then it would be a critical alert will be generated and then I wanted to bring in some metrics and the only thing I had to do here was use another out of the box check and just simply had to give the the value of the name this is a a running process that I'm looking for this will bring back a couple of different metrics with regards to that running process there is one more and this was also just for metrics and happens to be Falcon d so Falcon D is the Damon and Falcon sensor is a child process on the machine okay so that's the basics of how to set up the checks and the policies and now I would like to show you how to make it available the output available in the interface so I need to make something called a metric view configuration okay so this is my metrics view configuration and if I look at the metric types these are the metric types that we're picking up for this class called Cloud strike so CPU percent mem percent and if it's running or not those different metrics an additional step is to set up a metric configuration rule so I can tell the layouts platform what to do with this data so again we will look for crowdstrike just a very simple configuration we're just saying that when the configuration class is crowdstrike then I want you to do some data processing with these metric type IDs see if you percent members and running and in this case I have set the anomaly detection Action level to it alert so if an anomaly is detected it will create an I.T alert for me I could have chosen to just create anomaly alerts or just the scores or just the bounds or tell the platform to just take in the metrics only now that's the basic setup on how to get the data in um I'm going to do an additional couple of steps here because ultimately I would like to do an automatic remediation so I go to my event rules so this is for event management I have this very simple event rule for crowdstrike so the filter very simple the check name is OS dot Linux checkprocess and policy name is Linux crowdstrike then we've got a match and then I want to transform and compose the alert so I've added a manual attribute here called remediation action resource and I've given the value of Falcon sensor so that attribute will ultimately show up in additional information it's part of the alert okay so now it's on to alert management rules most of these alert management rules you see here are out of the box with the August 2022 release of the event management connector uh the ones that don't start with ACC are working agentlessly the ones that start with ACC are ones that I've created myself so let's dig into ACC start Linux service this is the one we're going to be using this is the flow that will ultimately get get kicked off when this alert happens the flow is very simple and the only thing I did was just copied one of the out of the box flows from the event management connector and did a very minor change what you see here is the input to the subflow this is really information about the alert so the conditions are matched in the alert management filter and then this flow is kicked off so we're just looking to see if a resource is empty and then ultimately this is the command that we're going to run on agent client collector so the important part is how do we know which command to run so what we're doing in this case is we're doing some very simple parsing so the first thing we want to do is create this attribute that is from the Glide record that is the alert itself and we're interested in a specific attribute called additional info additional info happens to be and Json so we'll take that attribute and that value and parse it into a Json object and then we are going to return a string so in this case since the objective is to start the service the string we will return is pseudo system control start and then this is where we're using that parameter that we created in a previous step which is remediation action resource so I go to my dashboard okay crowdstrike service it's red drill into that open the critical alert so this alert is on the crowdstrike CI and since we did the configuration earlier we can take a look at the metrics so these are the metrics specifically for the crowdstrike agent that are running on that machine we could also take a look at the Metro Explorer there's other CIS running on that machine like the Linux server and an Apache web server so we could look for things like Apache busy workers and then maybe CPU Steel just an example of some of the metrics that we can pull back but the primary focus was how do we get metrics specific to a brand new onboarded application into the IELTS platform so I've got a few playbooks here these were these are showing up here because of the alert management rule so I could create an incident or do any number of things really one thing I'd like to do is just so that we can run something like the next top via the agent so that's executed go back to the details here and we can see that top was run okay and in this case what I want to do is ultimately just start that service that's been stopped we could fully automate this process but in this case I want to be able to start it manually so I'll click this and that is going to start the service that was stopped in addition we have logs that have been taken in this has been taken in Via syslog so some of the accidents that are happening with regards to the crowdstrike agent are being logged in syslog so we're just picking those up automatically and in this case what I wanted to do is just do it very simple we're saying query click on this machine with anything regards to Falcon sensor in the logs let's take the last 12 hours and that will bring back a number of different matches so if an anomaly is detected with these logs ultimately we will get an anomaly alert so back to the um back to the dashboard the Ola is no longer in progress so it's it's no longer showing up here so we beat the clock on that one just want to show you real quick how we can easily make these dashboards so I'm in the platform analytics workspace just made a very simple custom dashboard and what I've done is just added a few new elements with the data visualization in this case this is for metrics so we'll take a look here configure and then we can see the data source so I can go into data source and for these visualizations we can choose tables indicators and look in the metric space or save searches from health log analytics or report on components that health log analytics is aware of so what we did in this case was we just used see the metric base to pull out metrics very specific to to um to crowdstrike so name starts with Falcon is the condition when the configuration name is it starts with vodka actually the name yeah okay so cancel that discard any changes so very simple to use machine learning to onboard new applications into the iOS platform very simple configuration to start getting the metrics and events and logs in also thank you very much for watching and I hope you enjoyed the show
https://www.youtube.com/watch?v=EnSYsy0G5uQ