logo

NJP

Public Cloud Monitoring with ACC

Import · May 28, 2022 · video

hi this video is about the new monitoring checks that we added to agent client collector recently these checks enable collection of metrics and events from public cloud infrastructure and services when running your workloads on public cloud environments like aws or azure you cannot always install agents on all your hosts and even if you do it your ec2 instances for example are not necessarily the hosts that are running your applications and databases so we need to provide an agentless solution using the agent client collector the way to do that is to use the agent in proxy mode as i will explain in a couple of minutes our new proxy checks for events and metrics include as per may 2022 storylist ec2 and azure vm as well as storage service metrics like ebs s3 azure blob storage container service metrics like amazon ecs and azure aks cloud databases services like dynamodb cosmos db and others and in addition we are continuously developing new checks for other public cloud services for monitoring on-prem hosts and applications we usually install agents on the host we plan to monitor but for public cloud monitoring we use the proxy agent configuration it means that we install an agent or any machine whether it's an on-prem server or a cloud vm then the agent will run this on this machine and use the proxy checks to run on the host in order to collect required information from the remote public cloud workloads such as vm instance cloud database or any other cloud service since we do not install the agents on the host we monitor we will need to use cloud discovery to create the cis for the workloads we plan to monitor the basic agent discovery will not be helpful in this case now let's move to the demo to see how cloud monitoring with acc works let's start with acc policies and look for the aws policies we can see here four new policies for aws by the way i already activated all of them because i wanted to show you some events and metrics so let's start with aws metrics first of all we can see here multiple ci types that are supported by this policy by different monitoring checks so we support a ec2 metrics ebs s3 rds and so on let's edit the policy in sandbox so we will better see the configuration so first of all monitored ci so the main ci is the aws data center i can also fine tune it with the regions i plan to monitor so in my case my workloads are on this data center and on two regions used to s2 and use is 2. by the way even though the policy filter is the data center and the binding of the metrics will be per ci type so for the proxy setting we see here uh that we use a proxy agent and a check so we see here multiple check instances for each of the services i mentioned for ec2 events we have a separate check again the policy filter is very similar data center and regions and it runs three check instances for the events dynamodb metrics again very similar policy filter proxy and one check instance we also have new azure policies i will not cover this in the demo but these are the new policies now let's go and see some alerts in operator workspace let's go to alerts and choose open alerts let me look for a group of alerts for example this one so we have here a group of three alerts first of all the ci type is aws data center usb2 and we have three alerts here three different alerts that are generated by each one of the event checks if we go back to the ace ec2 event policy we can see here three different check instances one for the cpu balance one filter and one for the network so you can see here exactly uh three alerts that have generated each one of them by a different event check let's open for example this alert so we can see that we have a cpu balance critical alert on a specific ec2 instance and while the other 14 instances are in okay state now let's see some public cloud collected metrics so in operator workspace i can go to cmdb then choose the ci and see its metrics i have created a list of predefined cis so let's choose for example one ci for ec2 instance so we can see that the ci type is virtual machine instance and here we can see the list of featured metrics or the metric tab that i can see each time i get an alert on this specific ci this is something that the user can of course customize but out of the box we deliver these charts let's choose the cpu utilization chart and open it in metric explorer and so we can add some more metrics from the full list of metrics or from featured metrics and we can play with the time range or we can choose last 24 hours or a custom range let's choose like the week last week and and we can also show alerts on the timeline so here i had an alert on this ci and some more metrics so same for evs let's use a specific ci and the ci type is storage volume and again metrics are collected we have featured metrics we can open it again in metric explorer so same goes to rds elb s3 and so on and with dynamodb i want to go back to alerts and to see the metrics from a different direction so let's go back to all alerts and search for a specific alert that is related to a dynamodbci the alert description says that the latency metric exceeded the critical threshold of 4 milliseconds now where did that came from before this demo i created a static threshold definition of a latency metric for dynamodb with a critical threshold of 4 milliseconds so what we see here is an alert that is generated using static threshold and not using the event check now we can see the metric tab or the featured metrics again in the context of the alert and not from the ci itself that's it for the public cloud monitoring using acc demo thank you for watching it

View original source

https://www.youtube.com/watch?v=eJMxmhXgn_w