logo

NJP

Beers With Cloud Engineers - Episode 17 - Service Graph Connector Update, SG-OTel, SG-AWS 2.X

Import · Aug 31, 2023 · video

well welcome everybody to another beers with Engineers session um so uh we'll kind of go through our standard Preamble before we kind of get into the the meat of the agenda for today so as always uh we we sometimes talk about forward-looking uh features on road maps and things that are coming but not fully baked into the product yet so always take that into account and don't make purchasing decisions based on any um any roadmap or forward-looking items that we mentioned our informal agenda as usual we'll just introduce ourselves talk about why we get together uh in this forum every month uh we'll do some uh we'll dive into a couple technical topics and as always uh you don't have to wait till the end for Q a feel free to come off mute or put questions out there in the chat or the QA whenever something comes up happy to explore uh tangential queries at any time why are we here so um in the context of beers with Engineers the answer to this existential question is that our goal when when Mike and I started having these little get-gatherers was just to have some kind of a forum where servicenow employees and customers who are exploring the cloud native world and and um somewhere on the journey towards transforming their their digital infrastructure into that cloud native landscape to talk about the challenges and discoveries and and innovations that they've been able to um to create and you know all the all the while trying to you know cast that against the backdrop of the servicenow platform um to try and show different ways that the platform can make that Journey easier and um and in many cases you know accelerate it so some intros so Mike if you would like to uh let us know who you are for sure so Mike Gallagher I am the manager of the Enterprise application platforms team here at drw um previously a servicenow employee um I have been in tech for longer than lots of folks have been alive um and I uh I'm just generally an all-around nerd anyway um I absolutely love all things kubernetes I'm a big proponent of it operations management and I love to basically Tinker with anything and everything that I can get my hands on specifically in the cloud native space um I um actually play a lot of board games and some video games and I would be lying if I said that I hadn't lost quite a bit of sleep to Baldur's Gate 3 recently so and today I am actually going to change it up quite a bit I am having a an ace guava cider it's a craft cider so we'll give that a shot well hi I'm will Hallam I work at servicenow as an advisory solution architect focusing in the itom areas I've been working in technology for quite a while lately I've been uh focusing a lot on Automation and specifically in in various Cloud native areas kubernetes serverless stuff the various public Cloud providers that kind of stuff really enjoy and find it gratifying to find ways to automate repetitive tasks I've never been a fan of doing the same repetitive mundane task more than a couple times and uh in my spare time I'd just like to hang with my family A place to pick up hockey and and probably too many video games and then I then I rightly should today I'll be enjoying a fiddlehead IPA from fiddlehead brewing in Shelburne Vermont nice okay so because I want to make sure I didn't forget to bring it up um before we kind of dive into our technical topics I did want to bring up the fact that we are uh I know I'm uh currently on the hook I'm gonna definitely be at kubecon I've made my uh travel Arrangements Mike I'm not sure where you're I know that's kind of your backyard from an employment perspective but I don't know how solid you are for definitely being there but um we're looking at doing uh for those of you who attended our live event and knowledge this year we're looking to do something similar to that at kubecon which is in Chicago uh November 6th through the 9th um still kind of early discussions as far as what form that will take so the topics TBD um just really interesting well you know kind of wanted to throw it out to this community let you know that that was in the works Garner some feedback if there was if if you're planning on going to kubecon or if you've got you know um peer organizations who are planning to attend as well um and just say you know do let us know if um if that would be of interest if you're going to be there so that we can kind of feed that back to the folks that are putting the planning together for the servicenow presence at qcon and use that to kind of shape what form the um you know the beers with Engineers contribution to our our presence at Q context I'm excited about that and I I likely will be able to attend but I have not finalized that yet okay cool yeah if nothing else I'm gonna be on the hook for you know serves now is going to have a a booth and I'm on the hook to spend some time in the booth talking kind of focusing on that um kubernetes observability and I Tom so like AI Ops event management uh Better Together story so that's kind of what I'm definitely going to be doing and then the more kind of beers with Engineers Focus pieces um still getting kind of discussed okay uh Tech Deep dive so um what we're going to touch on today is a couple updates in the service connector space that we touched on in previous sessions uh the first of those is going to be the um actually the topic that we explored with um Clay Smith from the product side at knowledge and um that was a really great time unfortunately the logistics didn't work out really well to like do any kind of recording so for those folks who weren't able to go to knowledge23 um they may or may not have had uh much exposure to this new service graph connector up till now so we wanted to touch on now that it's GA it's actually at a version 1.2 version 1.3 is coming soon and that's part of um and part of that process is the the product team is looking for feedback from customers after they kind of Kick the tires on the current iteration to kind of help guide those last sets of features that they're working to put into 1.3 so that's it's definitely kind of a very open dialogue right now on the product side with what the next you know iterations of the service graph connector are going to look like so if you've got particular feedback in that area definitely funnel that back you can send you know as always communicate that back through this forum um or you know certainly with your you know with your servicenow account team um so a couple slides to show the highlights of what's in the current version um it is it leverages um servicenow Cloud observability which was brought into our family by way of the purchase of lightstep um but this functionality does not require any kind of separate entitlement to the observability product the only caveat there is if you want to integrate with event management and set up thresholds within Cloud observability you do need a separate Cloud observability entitlement for that but to get the functionality that the service graph connector provides in the form of ready-made service Maps if you've got um open Telemetry instrumentation either explicit or Auto or Auto instrumented as well as a pretty turn key mechanism for discovering just your kubernetes clusters they so if you're an itom customer I tomvis you can use this capability without having to have any separate entitlement in the observability space and and effectively the way to think about it is the capabilities that you get out of it are aligned with the license that you have today right so if you have just visibility then you're only going to get visibility into the kubernetes Clusters if you have health you can get the Health Data but will also require the servicenow observability cloud observability license as well on top of that so um I mentioned TurnKey installation so the form that takes is uh there are already made Helm charts which can deploy this capability into a kubernetes cluster with some reasonable default values it's also got in in kind of it's a theme with open Telemetry where it's and kubernetes where it is very customizable so it's kind of a fit if you if you're Green Field and you don't have any open Telemetry in your environment today there's um reasonable default setups that you can deploy to those clusters pretty readily but if you're already you know on a path to using open Telemetry you can also just kind of tweak those existing open Telemetry deployments to then just be essentially just Fork off a p uh a stream of that data into the servicenow cloud observability back end to generate the needed metrics which are then imported by the service graph connector to generate your um to populate your cmdb generate your service Maps I actually when I was doing my initial testing of this I took that Helm chart and hacked it up and deployed it in a full-blown like get Ops model and was able to get it working pretty well that way so um I was actually actually pretty impressed with it sorry continue so without any instrumentation at all if you just install this in a in a an empty pristine kubernetes cluster what it'll do is it'll collect cluster node pod service and workload information um it scrapes those out in the form of metrics which then get sent to that cloud observability back end and then the service graph connector piece has import jobs on the servicenow instance side which import those in the form of corresponding Ci's so you end up with uh you know a cmdb populated with those basic kubernetes components then if you've got uh if you if you have workloads that are running on things like java.net python you can pretty readily activate open Telemetry Auto instrumentation and what that means is basically the open Telemetry collector can retrieve not just those metrics but also traces from those workloads without having to make any changes to to that code you basically add an annotation to the um to the kubernetes Manifest and then the open Telemetry collector will extract those traces without having made any changes to those workloads at all only in specific supportive languages though yeah it's java.net Python and I think there's one other one I can't remember the top of my head isn't it wasn't it go I can't remember it might be Go I mean that would make sense because goes a very kind of staple language in the kubernetes space and the open Telemetry space but we can we can confirm that yeah yeah and actually um one of the one of the links that's included in the Our Deck that we'll send out at as part of our follow-up emails there was a really great uh lab that was run by one of our observability smes at knowledge and he put the collateral the the materials from that lab into a public git repo which basically walks you through creating a an example environment within a kubernetes cluster that includes Auto instrumentation as well as an explicitly instrumented uh example application and so I'll provide the link to that public GitHub repo I found that to be really useful in my own you know testing environment that I ran with this uh service graph connector foreign addition to Auto instrumentation if you've got applications which are explicitly you know include those um open Telemetry instrumentation calls already then certainly you can tap into those that's kind of the ultimate end goal on a road map of implementing open Telemetry because that's what really gives you that very granular insight into what's going on and I'll kind of show you um in a couple slides what that looks like on the servicenow side and show it side by side with what you're seeing real time within the cloud observability space and it really provides a very detailed service map on the servicenow side to really kind of show you what those interactions and relationships are between various microservices so there is a new release like I said they're working on version 1.3 they're targeting end of year um one of the main features they're adding to that is to expand the CI classes that are captured by the service graph connector and what I mean by that is they're looking at uh things like um Cloud VMS so if you're running a kubernetes cluster the goal is to not just show you the kubernetes nodes but to actually take that one step further and show you the underlying um you know if you're in AWS show you the underlying ec2 instances that are running and kind of close the loop with regard to that overall picture of what infrastructure is hosting these workloads and the product team is still soliciting for uh customer feedback there so if you've got particular um CI classes that you would be interested in seeing come through as part of a you know discovering a containerized workload uh definitely funnel that back through the channel of your choice so this is uh this is a slide we brought over from our our time we spent going over this at knowledge and it really is just meant to capture where the pieces how the pieces fit together so on the left you've got the kind of the three types of of open Telemetry data metrics logs and traces funneling through open Telemetry infrastructure ending up on the cloud observability back end and that's providing kind of the ability to near real time interrogate what's going on Within These workloads and then from there you can establish thresholds if you if you've got a full-blown Cloud observability entitlement that's what lets you set thresholds so you can say if latency goes above this many milliseconds for example then you can make a call back into your standard servicenow event management setup to then trigger alerts which are then correlated with other sources as per kind of standard um AI Ops and event management functionality that you'll find within the platform and then the other piece down below is where the service graph connector comes in feeding the metadata into the cmdb so that when you get one of those alerts you're not just seeing Oh this particular container is or this particular microservice is having an issue you can see the full business context within the service map on your servicenow instance and see all the way up to you know here's the the business owner the business leader who's going to be impacted if we don't get this situation resolved before it causes some uh you know some negative customer interactions one quick key thing to point out as we're talking about like maintaining control of access of the data and where things are going that cloud observability platform is actually exists outside of the servicenow data centers and actually lives on kind of a custom like gcp instance um so someone's going to be aware of if that's a concern for you just understand that that's where the architecture lives for lots of customers it's not a problem um for us it would be a problem so it's you know yeah that I had kind of forgotten about that that uh yeah that's a great call out um so the cloud observability platform does still live in gcp um and that that does present a barrier to entry for some so it is something to be aware of when we're talking about this capability and this service graph connector yep um so this slide is just meant to capture the general kind of the way it's constructed and kind of the high points so as I mentioned there's Helm charts you can use to deploy this the the kubernetes side of the basically there's a there's actually a an open Telemetry operator which can take care of some of the heavy lifting with configuring the collectors for you um make sure they're running the latest version at any time and when you update parameters it automatically refreshes the pods so it it's a great way to kind of have more of a set it and forget it set up especially if you're in kind of that that more Green Field environment but it is optional you can also just roll your own open Telemetry collector configs and um use that as the mechanism for sending the the streams into cloud observability and then you know we'll we'll kind of we'll touch on when we get into the platform we'll we'll touch on the uh you know you basically got import jobs that are running transforms and making API calls into the cloud of cloud observability back end to pull in the data about the different CIS and also capturing those service to service relationships and that's um that's kind of where there's that's a big differentiator where uh with traditional kubernetes Discovery you were mostly going to be relying on tag-based service maps to draw those connections between different kubernetes components and this capability gives you uh additional Insight without having to collect any new data because of the fact is kind of sitting on top of the existing open Telemetry instrumentation so Mike if I remember correctly this was when you were kind of looking at this capability from a customer perspective these were kind of some points and questions that you came up with from the customer perspective that you thought would be kind of key to touch on as somebody is looking at this capability and then marrying it up against where they were with regard to kubernetes Cloud native within their environment yeah it's really you know a combination of that and like some of the thoughts and concerns right so um like when you when you stop and think about how you're using the data that's populating the cmdb today this is really thinking about the open Telemetry components and this service graph connector and how does it kind of fit into the broader picture so you know the first thing the first quick easy question to ask is do we have open Telemetry already either deployed or on the roadmap right because if that's the if the answer to that is yes then there's definitely um some you know quick easy ways that you can kind of show some value get this thing deployed and start pulling in that data from open Telemetry um and and then you know the real question from the real answer from that is okay where is that quick win is there a a particular you know service is there a particular cost for that's um already instrumented properly that we can kind of go and do you know like a quick POC on and and get that value showed right um the other thing is if you're using other service graph connectors already today right um the there would potentially be some overlap and kind of the methodology and the understanding and the fact that those are already you know coming into the ire and may also be able to like help populate and and broaden the data that you're getting across Hotel versus um other sources right the other thing to think about is like do you have somewhere else out there where you're building out maps of your Cloud native environments right um some cases that's a yes some cases that's a no and the like the capabilities can be wildly and dramatically different across different solutions right so being able to kind of populate all that data into the service now cmdb and getting it done based on actual real utilization and traffic across those components can be really really powerful right and this gives you data around like hey um you know what what's the current state what's the current sort of visibility and capability level um across across your environment uh one of the other key questions to ask is if we're already ingesting kubernetes CIS from other different methods how is this going to reconcile down with those other methods should we be turning off those other methods and just in favor of Hotel um or um or you know should we kind of tweak into the rules to ensure that we're augmenting that data set all across the board and then kind of last but not least is there any way in place today to be able to determine hey here are the components that are running in a cluster and and its infrastructure that's underlying underneath of it right that's definitely something that this can help with that a lot of a lot of organizations struggle with all day every day you know we there's always that sort of division between you know like the devs and the you know the infrastructure operators and the devs don't just care about what's going on in the pods and the services they don't care about anything else right so this kind of like ability to have that full-scale 360 view um is really powerful yeah so those are the kind of key common questions to stop and think about all right so um just to give you kind of a glimpse into what this service graph connector looks like when you've employed it in your environment um so here's an example of how you can deploy this pretty readily again so I did this with uh a couple kubernetes clusters one that was kind of just pristine with nothing running in it so you can kind of see what you get even if you've got like zero open Telemetry instrumentation anywhere and then I did it in a cluster that actually had an example open fairly robustly instrumented open Telemetry example app deployed to it so you can kind of see the difference so um this is that GitHub repo I I mentioned it's really useful because it's got just kind of a step-by-step set of instructions right down to the um the helm commands that you would run to deploy this into a cluster um I guess for the lab they use digitalocean to provision the Clusters themselves but I used Amazon eks and didn't have any any issues so seems to be pretty agnostic as far as that's concerned um so yeah I mean just you know to really quick touch on what it takes there's this set of Helm commands here basically to pull in some repos it does rely on the cert manager um Helm chart to manage the TLs certs that it uses to talk back and forth with the observability back end um it seems like this is pretty commonly in use out in the field anyway in fact our new automated uh kubernetes certificate management offering um also makes use of the the cert manager uh Helm charts and then you just install this open Telemetry operator again it's not required it just makes things a lot easier operators and kubernetes are they're basically kind of like wrappers that provide automation around lower level you know deploying lower level um workloads and so what that looks like on this uh when I deployed this against a very generic cluster is I get this CI and related Ci's in the form of I've got my list of PODS my list of services namespaces and then various you know um payload or workload types deployments replica sets stateful Daemon sets and I basically got this all just by deploying those Helm charts and then running through the guided setup for the uh service graph connector itself which is is fairly minimal essentially the main thing you have to do is um and you know does rely on the cloud observability product so if you want to look at this you can either go self-register for uh an organization within the the light step platform or your servicenow accounting can facilitate that for you and basically just give you you know basically set it up for you and give you the you know the login credentials that you need to get into that so once you've got that organization set up that'll come with an API key one or you know one or more API Keys you plug that into the service graph connector configuration and it's a standard you know it's a standard credential type record um and then once you do that you just uh go to the next step to test the connection and when that connection tests successfully it allows you to pull in the kind of the um the container that cloud observability uses to collect pieces of a workload is called a project and so what that does is once you've got um you know a successful connection the guided setup will go through and connect to your cloud observability organization and pull down a list of all the projects that are currently there and that's what we've got here is a list of all the projects that are in my test organization on the cloud observability space and and so um after that set up it's really just a function of scheduling the import jobs to run on the Cadence that you want and then those just kick off and pull down updates from the uh cloud observability back end based on the frequency which is desired and so what that looks like when you've got a fully instrumented open Telemetry compatible application is something like this so this is the the light staff formerly called lightstep now called servicenow Cloud observability UI and what you're seeing here is the service diagram that the observability platform creates in real time and it'll kind of regenerate that dynamically this happens to be a latency histogram so if I want to see the services that are involved for kind of my most um my highest latency types of interactions it'll generate that diagram kind of on the Fly and show me the different services that are interacting and what the service graph connector does for you is it basically produces a very similar service map to that Dynamic service map that you get on the observability side but it lives within your cmdb so if my product catalog service is experiencing an issue I can see all of the upstream and downstream impacts that that could have within the servicenow platform okay I always forget to hit the uh start a current slide right when I restart the presentation back up while he's getting back to that slide any questions or thoughts on the hotel piece before we shift gears don't be shy so this QR code that you'll see in the deck that's the quick link to the store Page um if you would like to uh to download it and try it and and as I mentioned you know there is that there is that kind of self-provisioning piece on the observability side that account teams are totally on board to help with so don't let that be a barrier entry if it is something you'd like to look at they can help you to get that uh observability organization set up with uh without any uh overhead on your part so the the second big service graph connector update that happened over the summer and had a bunch of new stuff that it brought to the table was the um the 2.0 for the AWS service graph connector and we had um we had merly from the product team come and kind of give us a a preview of what was coming with that um and now that it's actually out it's actually up to version 2.2 uh and they've added some even cooler stuff like eks support which is uh which I think is really cool because when um also in the same time frame we released the the 1.0 of the gcp service graph connector and when I was playing around with that one thing that struck me was hey gcp actually exposes kubernetes cluster information via their native Cloud API so just by doing that service graph connector you could all of a sudden get a lot of kubernetes Ci's into your cmdb without having to ever set up any separate kubernetes Discovery I thought that was really cool um I suspect other customers or customers thought it was cool and asked hey can we do this with AWS and unfortunately uh there was a little more work involved there because AWS doesn't expose that at the API level like the other two big players do because it looks like Azure does as well like I can go into the Azure portal and see kubernetes namespaces and pods and stuff um so anyway um with version 2.2.0 we now have the capability to look into eks clusters with just the service graph connector and there is some setup involved and we'll get into that in a minute um a big plus for this new iteration is multi-instance support so that's not only can you do you it's it's kind of like sky's the limit you can slice and dice it however you want now previously you had to basically start your AWS Discovery via service graph connector at your top level org your AWS organization you had to have some access to your master AWS account um and if you didn't have that the you couldn't you couldn't set this up not only is that no longer a requirement but if you've got multiple AWS organizations you could actually set up multiple instances of the service graph connector within your instance now to walk across multiple organizations uh as I kind of touched on you can also install this to a single account and speaking selfishly as a solution consultant this was huge for me because up until this version came out I couldn't even demo this tool for customers because I don't have any access to the servicenow main billing AWS account so I could not get this thing set up in my demo instances from a customer perspective it makes piloting and kicking the tires on this infinitely easier because all you need is just uh a sub you know a regular sub account within AWS you don't need any access to that AWS organization in order to test and validate all of the functionality that this service graph connector offers automated key rotation super handy um one of the kind of one of the hitches that you could run into from time to time when it when you're talking about the service graph connector is the fact that unlike pattern based Discovery in the cloud you couldn't just use like a mid-server role directly you had to have that initial uh API key to let you to kind of establish that initial connection and then you could assume roles from there to hop from account to account but you still needed that kind of that bootstrap credential and in this day and age the idea of you know kind of static API keys and secrets is a little anachronistic it's not generally what people would prefer to use when it comes to you know API access to discover their Cloud accounts um the automated key rotation I found it works really great and it mitigate that mitigates the use of those kind of static credentials because in my test environment I just set it to a one day life and it just rotates the credential every day no problem and then eks support again that was I I saw um in some of the community articles that merly put out there where he really got in deep about the different configuration parameters for this service graph connector because it is very flexible especially now um there were comments saying hey can we get eks support can we get AKs support and and now that's there and it really um I I think I I think on the one hand we've got a lot of options now to discover kubernetes right um on the other hand if you're just starting down the journey and you're going to set up a service graph connector anyway the fact that you can now pull in at least your basic kubernetes objects without having to go through any other setup for any other capability I think is is a big plus especially given the fact that you know gcp it's there with gcp today and um I it didn't look like the Azure service graph connector did it yet but Azure definitely makes those details available via apis so adding that to the Azure service graph connector is uh a fairly straightforward exercise because all the metadata is already available via Cloud API calls yeah yeah so just touching on some of the benefits um yeah multi-instance support it's just uh We've Got Big customers out there who have multiple orgs you know they'll have like a prod org non-prod org you know gov Cloud org non-gov Cloud you name it so now that's no longer a barrier there's a separate set of guided setup that lets you add additional instances whether it's Standalone accounts or separate AWS organizations installing into a single account great for Pilots great for kind of evals it makes that whole process a lot more agile especially when it comes to debugging AWS permissions which I I don't think I've ever deployed anything to AWS where I didn't have to kind of poke around with permissions to some degree to get it to work exactly right so the ability to just deploy that all self-contained into a single account before you have to worry about stack sets and deploying it you know organization wide it can be a real big streamlining Factor again great that now your solution Consultants can demo and test it for you much more easily is awesome uh lower complexity and Implement on a small scale if you just have some targeted deployments that you need to run this in you don't have to boil the ocean anymore and handle it at a AWS org level yeah uh the automated key rotation it is optional you turn it on you turn it off I found that um I did have to tweak the permissions a little bit because um and I'm not sure if it's a bug or a feature yet but so you've got your basic like your service account that's what has the API credentials and then the service account has the ability to assume a role and the role is what generally has the bulk of the permissions that are required to do the discovery um maybe because it's a fairly new feature what I noticed was when it's running the key rotation it's not assuming the role it's running the key rotation um commands API calls as the service account directly so I had to add the create access key delete access key permission to the service account directly as opposed to it's already actually in the role because we provide cloud formation templates to create the um to create the role for Discovery and that's already got create access key delete access key I had to add it to the user in order for it to work so that may be in a rata that's just going to get fixed but something to be aware of if it is something you want to take advantage of immediately today uh the user configured frequency works great it defaults to I think 90 days I just crank it up to one day because I was testing it I wanted to see it work and I tested I checked today and my API Keys like 12 hours old as of as of now it's a scheduled job so obviously you can run it on whatever Cadence you want um I think the granularity for the property the the the property of how often to change it I think it's in days though so it does make a lot of sense to run it more often than than once a day the eks support is um it piggybacks off of the existing systems manager functionality which has been in the service graph connector for a while now so that uses systems manager to gather things like the software inventory as well as to run some local commands so that you can get things like the serial number the OS level serial number from ec2 instances without having a mid server in the environment which would normally you know SSH or Powershell into those boxes and so you do need that that's an optional feature like you don't need to have that set up to run the service graph connector but if you want the eks support you do have to have that SSM stuff set up which is basically you're giving solution um you're giving systems manager permission to run some specific commands and then you set up an S3 bucket to store the output from those commands because the output can be uh larger than it can return back in a single kind of in an interactive uh API call so what it does is it stores that output in an S3 bucket and then the service graph connector retrieves the output from the S3 bucket so it uses a similar functionality you just add a couple additional commands to what that systems manager agent can run and it just uses um it basically uses Cube cuddle on a Bastion host to retrieve the information from the various eks clusters so you heard me say Bastion host you do need um you do need at least one ec2 instance that has the ability to connect to those control planes that the eks Clusters are standing up so um you know worst case if you're a very segmented Network and you don't have a lot of uh you don't have like a a Transit Hub that facilitates communication to those different control planes it could end up you need a lot of bastions um I would say if you're looking at that kind of situation it would just it would definitely warrant trying to maybe look into something like a Transit Hub or something um to kind of make those endpoints more centrally accessible and then limit the number of of bastions that you need to provision that's definitely kind of a feature where um you know it's not very attractive probably to have a whole bunch of separate bastions but the solution itself does support that it lets you create records for each Bastion and it kind of maps them to each cluster so that it can it can handle many-to-many relationship without a problem it's really more on the kind of the the setup side that you'd probably want to minimize the number of Bastion hosts that you need to allow this eks support capability to function and uh there's uh dedicated knowledge base article for the eks support there's a link included in the deck that we'll send out with all our slides because that's really key in using to to get the things set up I did run into a bug uh checked with support and there's an active prb record for it and it's going to be fixed in the upcoming version 2.2.1 has to do with when it pulls in namespaces for whatever reason the namespace records come in and they don't have any name for the namespace so you see a bunch of dependencies on a namespace record but it doesn't tell you what the namespace record the name the name of the namespace is um also there was a script include that I noticed um whoops uh there was a script include that is responsible for downloading the output from S3 and it so basically what happens is it runs the job to kind of run the commands and dump the output in S3 bucket and that's asynchronous so then what it's supposed to do is it's supposed to check the S3 bucket and then just kind of loop until it sees that output file show up in the S3 bucket the problem is this this script include that I call out in this slide uh was checking it wasn't checking for a 404 I was checking for a 403. so the problem was it was getting a 404 and it was bypassing the retry and immediately just saying oh I can't get the output and it was and it was bombing out so I've got a case open on that as well um but in the interim it was pretty easy to to fix that it was like one just one line in a script include to to fix that and then get my output downloading successfully and so right now what it looks like is it's pulling cluster node pod deployment and service CIS but as I mentioned it's just pulling uh like yaml format or Json format payloads using Cube cuddle commands so it's pretty extensible in the short term and I'm sure in the long term it's they're gonna just keep adding additional CI classes as the product matures so just a couple looks at what that looks like on platform so the the guided setup as I mentioned here's like the the monster kind of one size fits all configuration page that we have now for this um service craft connector and it's basically intended to account for all the possible scenarios so you don't have to fill out every single one of these forms it's really dependent on what your specific setup looks like and there's also a detailed Community article about all of this there's a link to that on the Links Page that'll be going out in the deck so don't feel like you've got to kind of absorb all of this just wanted to give you a kind of a look to how it how it's set up now with this uh generation of of the tool so the two things you have to give it are connection alias and the name of the role that it's going to assume this section here I'm running in a standalone account but as part of the multi-instance support they added basically the the ability to tag a given instance with an organization you know some type of an organization I use my account ID and just made up a name so that when you run the diag tool that they give you you can select which um kind of which tenant which instance you're running against for it to do the validation and then yeah I mean it's really I I because I use the assistance manager setup because I was testing the eks functionality um I did fill in my S3 bucket these values down at the bottom uh come with defaults so I just left them and those are the names of the uh what are called the systems manager documents which contain all the shell commands that are executed for the um this is for the Linux and windows Discovery to pull things like the serial number information and then these are the two documents that contain the shell commands that are used to list all your eks clusters and then interrogate each eks cluster in turn to pull down basically this big dump of all of the metadata containing the CI classes that are currently imported by the service graph connector so here's an example of a ec2 instance that's inventoried by the service graph connector um and so because I've got the systems manager inventory enabled and I've got my systems manager run command capability setup it actually can pull down some of this related data that normally you'd need a mid-server for so you can see you've got the running processes listed as well as the the software inventory on this Linux server which is running on an ec2 instance and in terms of what a kubernetes cluster looks like let's see add up so this is my eks cluster and when it fleshes out the relations you'll see that you know it's filling in um it's filling in pods it's got my node information down here as well as service CIS uh one thing I did observe today is it looks like right now and I haven't I just kind of ran into this right before we uh we started up this session it looks like it might be combining stateful sets into the Daemon set um CI class so I'm going to follow up on that it looks like I suspect that's probably just um a bug which is going to be fixed with the the upcoming two two one because it's kind of like a one for one when I look at if I just look at stateful sets um I think it I think the table's like empty but if I look at my cluster I've got you know I've got staple sets in there so the fact that they're not coming in um but if I look at uh game insets if I remember correctly there was a whole there's a whole bunch of them but when I look on the cluster there's definitely a lot fewer showing up on my cluster than I'm seeing in cmdb so I suspect that they're just kind of conglomerating multiple workload types under damage set when it's um staple sets oh and replica sets that was the other one um I think most of what I've got is a replica sets because that's kind of the subordinate workload that whenever you do a deployment that gets created so I think if I look at the cluster if I look at replica sets it shows me a whole bunch but um within the cmdb everything seems to be right now it seems to be kind of merging at all sticky at all under Damon's set so that was a little bit of an Errata that I did notice um I don't anticipate that's going to be difficult to get remedied though uh another helpful link just kind of as a shout out uh you know as I mentioned with AWS a lot of times um permissions can be a bit of a bear so um I did include a helpful doc link to an easy way to query uh cloudtrail API logs which I found really useful when I was kind of walking through setting this up um the diac tool is great it'll tell you this permission failed but it won't tell you why or it won't tell you exactly like what role was being used verify that it assumed the role that kind of thing whereas when I was able to query the cloud trail logs and a lot of companies have you know user friendly ways to do that anyways but if not um I included a link that makes it kind of easy if they're getting dumped into S3 you can use Athena to query those logs and see show me access denied against S3 and it really lets you home in on where the permissions are are running into trouble so we traditionally run for an hour and we're at the top of the hour so I want to be respectful of people's time um I am not constrained um so I'm certainly available um Mike I'm not sure what your what your uh schedule's like today awesome but um yeah I think I don't think I had any other kind of uh points the obligatory points I wanted to make sure that I touched on so certainly happy to dig into any of the stuff we've we've talked about today further or completely go off on a tangent if folks have uh questions or topics that they wanna that they wanna air in this forum and we should probably conclude the recording at this point yep good call I will do that

View original source

https://www.youtube.com/watch?v=rj8OnVzzdtc