Beers With Cloud Engineers - Episode 6 - Managing Container Vulnerability with Prisma and ServiceNow
welcome everybody to session six of peers with engineers um looking forward to a great call today as usual um we're pretty relaxed so if you have any questions don't hesitate to come off mute and jump in and ask if for some reason we didn't get a chance to allow you to unmute yourself just raise your hand and we'll make sure we get that for you okay so as usual we'll start off with this safe harbor notice right some of the things we may discuss may be not released yet um maybe forward-looking things so uh don't make any financial decisions based on any of the conversations that we're having here today certainly not when it comes to stock so all right so well as usual again as you as as is normal we're pretty informal so um short sweet agenda as always um but let's always go back over why are we here and and what's the goal of the calls right the most important thing is uh there are a lot of our customers in the servicenow space who are moving into cloud native capabilities and containerization and both will and i kind of felt like there's there's not enough resources especially not enough community building in that space so we wanted to pull this together to give everybody a place to learn from each other learn from the stuff that will and i are doing and then talk and continue to grow so who we are my name is mike gallagher if you haven't met uh this is your first time uh welcome i really appreciate you coming and joining us today um as it says uh i've been in the it industry for a super long time and have if it has a one or a zero in it i've probably touched it or managed it in my career um most recently kubernetes has become a really big passion for me um so that's a big reason why i've been working with will to to run the sessions um but aside from that i love playing board games um i love um music we were actually just talking about going to shows before the the the session kicked off today and i trained brazilian jiu jitsu so i love that stuff hey will over to you yes sir uh will hallam uh itom architect i've been in i.t for a couple few decades i'm currently focusing on automation with a particular concentration in uh cloud which has uh more and more had its fingers it's its tentacles into the the kubernetes world it seems like it's it becoming more and more a ubiquitous topic so definitely spending a lot of time a lot of time on kubernetes related topics and my spare time i like hanging with my family and playing video games and pick up hockey and per are the title of this series i will be cracking open a long trail little anomaly ipa a delightful uh light low-calorie ipa which still retains a a nice hoppy taste which i appreciate and as usual i forgot to introduce and open my beer until after will reminds me to do his so i am actually have been so much in love with these admiral abyss that i'm gonna have another admiral abyss this month um it is a chocolate milk stout from odd 13 brewing which is a local brewery here in in colorado and they always have like the best art on their cans so super big fan all right hey i'll throw up a quick poll just to kind of get the check of temperature on the room um so the the brief poll just with a couple questions around container or container image scanning since that's a big topic that mike will be going into specifically how we integrate with a popular container vulnerability excuse me vulnerability management solution so if you have a few seconds to fill out that poll i'll just kind of leave it for a little bit back over here while we wait for that why don't we have our esteemed guest linda shinkle come off mute and introduce herself um and then talk about how we got here yeah thanks mike um linda schenkel i'm a platform architect with servicenow been in the it world for a comparable amount of time 20 plus years and actually live in colorado by mike so we've worked together for quite a few years i guess now it's been um but yeah i'm working with one of my customers and mike's my go-to right on anything cloudy to use your term and so we're got a customer that was asking for some help with assignment group relationships to vulnerabilities that are being related to containers they were using prisma as a vulnerability scanning tool yep there's the problem statement right there so you know the key thing um this was kind of an interesting thing and and and the reason why we're all here today is actually because um they came to linda and said hey here's here's this problem here's what we're trying to solve how how can we help do this and because linda is a platform architect that's assigned to work with this customer she's incredibly familiar with the platform and with their particular configurations and how they work and how they operate and so she reached out to me and said hey what do you think about this how can we get this working and i immediately said this sounds amazing i think this is going to be a really really good project for sort of end-to-end scope especially because it's so powerful from a use case perspective so i think let's let's dive in and see what we can do so um will has ended the poll let's go ahead and share the results all right so go ahead will um so it looks like about half and half in terms of um of the respondents 50 are planning to scan containers and or images but haven't launched any initiative yet and then the other fifty percent are actively in the middle of implementing uh and then in terms of overall approach 50 third-party tool 50 not currently investigating um so i guess that probably aligns to the 50 that have plans to do scanning but just haven't really set a direction or started uh implementing yet and then in terms of an approach looks like current plans aren't of the 50 who responded current plans in terms of tracking the you know actually tracking okay we know this vulnerability is out there but who's responsible for remediating them and when is that remediation completed that kind of thing um they're indicating that they're using some separate third party or or uh potentially an internally developed tool to track the actual work that needs to take place in order to resolve the vulnerabilities once they're identified makes sense so linda i'll let you kind of give the rundown on the solution and then i'll talk to like wind into the deep dive of kind of how it all went yeah sure so we started to look at the um at the data and found that we had container vulnerabilities with cvits coming in from that prisma integration and they were relating that vulnerability to the docker image ci um which at first glance my customer was confused about and thought that it should be maybe down into the instantiated containers and after talking with them really deciding that the owner of the vulnerability really should lie in that image that produced the the container because that's where the vulnerability came from um there there was a use case that we had previously to this one that introduced the common service data model technical service technical service offering and ci group um or assignment group synchronization that capability that came out quite a bit ago and so the my customers used to using that so what we've decided is that every docker image ci now has a new compliance or new requirement within the scene to be to own a technical service and technical service offering relationship and then we will use that relationship to assign the um cvit to the appropriate group to resolve perfect so that's uh a really good kind of overview of what it looks like but i want to take a step back and talk about how vulnerability response works on the servicenow platform and then i'll quickly jump into the um of the prisma integration what it does and then i'll kind of show the the end-to-end data flow and ownership so um the first thing the way the vulnerability response works is we have various different vulnerability scanners right in this particular case we're talking about prismacloud but it supports tons and tons of different stuff but it goes in and it creates those vulnerable items based on what it finds in its scans so in this particular case what linda was mentioning are those sieves the container vulnerable items that were created on the servicenow platform and then because we have all of this information in the cmdb and we know kind of how it's all interrelated we can prioritize that based on the business risk right business you know the threat and what the risk context are is for those applications and so then we can assign that out to the appropriate teams in order to be able to say hey this is your vulnerability you need to go and manage it you need to go and remediate it there's also some hey this isn't my thing we'll we'll send it back and reassign it to the right team and then um that has a whole trackable workflow as as is a standard with the servicenow platform and then once it's been actually remediated we now not only have the task um that says hey this has been completed and remediated but the scanner also can go through and say oh oh you know this this is no longer a vulnerable item so it can automatically close those for you as well so so the whole process is really sort of managed from an end-to-end perspective via the this secops vulnerability vulnerability response module and that's not something we normally talk about here but because this is such a powerful use case i wanted to bring up sort of how it works and how things work so let's talk about how the prismacloud vulnerability scanning works in particular right so um it's prismacloud if you've not heard of it it's a an industry standard vulnerability scanning toolset it has the ability to scan lots and lots and lots of different things for vulnerabilities things like cloud configuration things like your code repositories in this particular case we're talking about the container vulnerabilities today it's it's an absolute market leader kind of in this space part of the reason why i've been involved with them and heard of in the past was they acquired twist lock a couple years ago which was a container security platform that was very very useful and very powerful so in in this particular case in this customer's environment as part of every single build in their ci cd pipeline for container builds the the last step is hey this gets scanned by prismacloud compute and it looks for of any vulnerabilities in that container image and it can it has workflows and tie backs into the surface now platforms where it can like stop the build it can stop things from being provisioned based on that but that's not how they do it they just say hey here's the list of vulnerabilities and and then what's interesting is it does it at build time and in addition it does it every night it scans all of those repositories that they have already in place so if during the day or a couple days later now the vulnerability database has been updated and things that looked okay a couple days ago now all of a sudden is a vulnerability and we have to be aware of it and so at night it scans that repository it finds that vulnerable item and then it creates it that vulnerability record within servicenow for the users to manage and and remediate the entire process and so and this is where the whole process gets very convoluted if not if you're not kind of clear on how things are going to work and where the data all lies and what the relationships are which is why linda worked with with her customer to build out this design and understand kind of where the data flow is and what the relationships look like and this right here is is really the power of the platform and how it builds out those relationships so everything in purple is data that was built by the prisma integration right in this particular case this customer is using azure devops so that's the kind of lighter blue red is actually what's running in gke and then green is the data that's in the um in the cmdb right so prisma as it goes through and it builds those vulnerability findings right it scans the images and the blueprints inside of azure devops and it also ingests ownership and other kind of details around who's built that image and where it is and then that information comes into service now as vulnerability entries and vulnerability images and then there's these vulnerable items those are the cvits that we were talking about and those cvits have relationships to records in the cmdb so we're we're discovering all of the information in their gke clusters and this is something that we've we've worked on with them over the last probably six months to a year and we're discovering all of that information and so that's all being populated in the cmdb so we have the clusters that are there and those and those pods that are running inside of that cluster and those pods contain a docker image right and those docker images are are instantiations of a docker container right but in but ultimately the container repository has a repository entry that relates to the docker image right so all of these relationships are all there and available so when we go back to hey let's think about this from a reporting standpoint right i need to now find all of my vulnerable containers that might potentially have log4j just to use a recent painful example right i need to find all of the containers that could potentially have the log4j vulnerability and so we can go and look for those vulnerable items and then because we have relationships to all of the docker images from those vulnerable items then we can back from there into all of the running pods that have been instantiated from those docker images and then we can get to the owner we can get to the application we can get up to the business service that's consuming that application so now we can do things like assessing our risk to the business services and the business capabilities based on what the vulnerabilities are that are running inside of those containers right so um really really really powerful capability to be able to see all of that data from the the image that was built and what was vulnerable all the way out to hey this is actually running as multiple different pods right when you start to think about this if i have a a deployment that is running five copies of the exact same pod right then i now have five different containers and all five of those containers might have that exact same vulnerability so i need to understand like is my attack footprint this big right or is it massive right um so a huge benefit to be able to have all of this data have it all tied together and really has helped kind of solve a problem for this customer so a couple of things that i felt were important to call out uh as um things that we kind of bumped up against and um want to share with folks so they don't run into those same pitfalls so make sure that you're discovering your entire kubernetes infrastructure right because if i'm not finding all my running pods and all of my other components i won't have the necessary pods to be able to tie back to the docker containers and the docker running images so that's important in addition to that there's a there's a pattern specifically called collect container repository that is enabled by default but it's in some places i've seen it deactivated if that's not activated um it's a it's kind of a sub pattern that's called by the kubernetes patterns um to be able to get the container repository information about the docker images that the pods are running from instantiated from and so that process helps to build the necessary relationships that the cvits get tied to so it's important to ensure that that pattern is running and that it's going to work okay and that you'll get all the necessary data the other important piece and this is something that is really really important is the the in the ingestion of the data from the prismacloud integration uses a scripted rest api to push data from prismacloud into the servicenow platform and then once it's in the servicenow platform all the you know vulnerability response processes work perfectly fine but that scripted rest api has some configuration necessary in order to um ensure that the right matching is occurring so that we're not creating duplicate cis for these container images so some container images entries will have a prefix appended to them which could cause things to not line up and now all of a sudden instead of having all the right relationships you know down to the right ci's now we're having duplicates and so that could be a potential problem and something to be aware of and i provided a link to the documentation which actually is out on the palo alto website that talks about how that scripted rest api is configured and how it all works so questions about any of that and how that all works i don't have a running environment to actually show you guys um specifically because this particular customer preferred to remain anonymous and i didn't want to show their actual data so um i i wanted to open up for questions see if anybody had any thoughts i have a quick comment i i've noticed that behavior you just talked about with the um in the repo table or the image table where depending on how an image is discovered sometimes it's just like a shaw hash sometimes it's got like docker d colon or container d colon and then the sha hash um and yeah that's just kind of a recurring theme we've got uh there's a poc exercise in which i'm participating where this customer happens to run a lot of they're all on prem and so their regular discovery they they have kubernetes and raw docker and there because all of their stuff is running on their servers they are can actually discover the docker runtime kind of in parallel to the discovery that's happening from the kubernetes api piece and so i anticipate they're going to run into some normalization um opportunities to avoid having kind of duplicate image records one that comes from the docker engine directly and then one that comes from the the kubernetes because i definitely noticed between the two there's the the kubernetes one is the um that's generally a little more descriptive where it'll say like docker d colon or container d colon and then the hash whereas um docker basically says yeah docker is perfectly happy to give you like the little shorthand not even the full hash it just gives you like a little shorthand for the image um so there's depending on how heterogeneous uh the discovery sources are and the environment is you may have to make some adjustments there so that everything lines up and it's all normalized and you don't have duplicates but once you've got that dialed in the ability to map that all the way through right i go back and look at this the ability to map it all the way through from the vulnerability all the way down to how many containers do i have that are running this and what the business impact is is is really pretty powerful cool other questions so so would a best practice be to kind of be careful of the scope of your discovery if you have you know just plain old docker out there while you have kubernetes also so that you're you're kind of controlling what you're finding with the with the docker only pattern it would probably i i i definitely think that the more sources you have the more good sources you have for your cmdb the better off you are because it's less likely that you'll end up with gaps um but from an implementation standpoint i would definitely say don't turn everything on at once like turn on your kubernetes discovery make sure that's all stable and then if you do have you know docker runtimes that you're able to discover directly then there's definitely value to doing that um in fact your question is kind of a nice segue because one of the things that i have prepared to show today is a little exercise i went through where i can actually interrogate a docker runtime for what processes are running inside containers which is something unique you have to own the server in order to do that but we had a customer who asked for that and so um you know so that that's an example of where you know that docker pattern can add value on top of the the kubernetes pattern for the for the customer that happens to be you know running their own um kubernetes estate versus you know getting uh kubernetes as a service yeah yeah that's that's really you know vanilla on-prem kubernetes versus eks or aks or gk or something like that yeah and i thought the i thought the catch too was is that you're not just running kubernetes you have just standalone docker engines out there that are not part of a larger kubernetes installation yeah i mean that was the that's the scenario i'm dealing with um right now is a customer that's got it got docker being managed by rancher or sorry kubernetes being managed by rancher they also have a bunch of raw docker out there that they also want visibility into okay okay three questions any other questions thoughts concerns around how this all works okay cool with that i think we'll move on to the tech deep dive and will if you want to steal a share feel free alrighty so as i alluded to um what i have today is a kind of a a real life story about a scenario that i've run into as part of a poc exercise and this customer does have a diverse mix of containerized workloads and one of the types of workloads that they'd want visibility into is just raw docker containers and so as we're going as we've kind of kicked off our poc and we started walking through some basic discovery and after running through the discovery schedules they were clicking through different ci's and they clicked through a server and noticed that on uh you know and i'll just kind of do that just to make it visual if we pick uh you know for example a linux server those of us who uh are familiar with browsing through cis will find the uh you know the tabs that populate at the bottom of a ci record familiar and you know in addition to the exhaustive relationship information we've got these tabs down at the bottom one of which is running processes and and so then they moved on to a container ci which was being produced by the the docker pattern and they brought up the record and looked at it and said well this this is nice but uh we really would like to have that that running process tab as well so we can see what's running inside um inside the container and then potentially you know carry it further and do app discovery in addition to some of the other approaches that they're talking about which are you know pulling s-bombs software build materials from the images themselves which you know might kind of alluded to the fact we've got all of that information we've got every image that's in your environment that servicenow is discovered populating a container image table in your cmdb and so from there what we're looking at is to then go through and pull software build materials for each of those container images and kind of fully fleshing out the cmdb to embrace the containerized side of things to the same level that we've done with the server side of things but that's so that's another you know that's another aspect of the poc that we're looking into but their initial ask was well can we just get some running processes for our containers and when they asked i wasn't really sure that's not my head because that wasn't something that i had done routinely most of the container type work i do is in the kubernetes space and so i don't tend to spend a lot of time right down working directly with the container runtime but a quick perusal of the the docker documentation revealed that there is a docker uh sub command which is top that you can use against the given docker container to get a process listing and it takes all the same arguments as the unix ps command so that with that piece of the the puzzle kind of solved i then just looked at the out of the box docker pattern and saw that it was already just as part of what it already does it's making extensive use of the docker cli just executes various clis to pull back payloads having to do with what containers are running what images are on the system and then so then the the next challenge that i had was well how do i um how do i pull back the information that i'm looking for and kind of set it aside so then i can align it with the different containers that are running on the system uh if you're not familiar if you haven't played around with editing patterns one of the the basic aspects of a pattern is it wraps everything it doesn't have it doesn't really have the ability to do a discrete kind of for for loop where you say pull a bunch of stuff in and then iterate through it and um and do a particular action on it it kind of the way it approaches it is through various transforms and long story short it was it was not it was not straightforward from a drag and drop perspective to just say yeah loop through all these processes that i'm finding on this um i'm finding on this container and then just stick them in there with the rest of the container information that you're already pulling out of the box so i i did a little more research and came across this extremely useful feature of our discovery system which is called uploaded files and the uploaded files mechanism allows you to create any kind of a script that you can then that the discovery system will then upload to the the ci that's being discovered and allow you to execute and then produce a payload which you can then pull back right in line with your pattern so what i did was i just created a couple quick python scripts to enumerate information about my docker containers so the first one i created was i just called it docker ps and it just um it just does a docker ps command and pulls back each of the container ids and then it loops through those and runs docker top against it with a specific format pulling back the pid pid and the the full command string then it does a little more parsing of that and then provides that as a payload back to the pattern and then just a little preview i i kind of moved on once i got the process listing working and now i'm i'm working on a way to enumerate the software that's actually running the software package inventory that's running inside a given docker container and i'm just doing that right now using the experimental docker s-bom software build materials sub command that you can install as a plugin and that's uh it's pretty slick it actually it's based on a popular open source project called sift which is a very powerful image inventory tool which it literally just performs it's basically like running an rpm listing against a server except on a container image just dumps out a whole list of rpms or you know apt packages dev packages that are built into a given container image so once i have my scripts and uh tested them out to make sure they produced the output i was looking for i just kind of inserted them at applicable spots inside the uh inside my copy of the out of the box docker pattern so the pattern designer includes a put file operation so you basically just say put file you select a file from the list of uploadable files and then you assign the location that it gets stored in to a temporary variable in the pattern and then at the opportune time i invoke that script in this case i do so after the uh after the image uh metadata is pulled down i think that was kind of a little arbitrary i clicked for some if anybody knows a way to like pick up a pattern step and move it um feel free to let let me know either live or send me an email because um ideally i'd want to arrange these steps a little more logically but i have yet to find a way to like you know you think you can kind of drag and drop these steps kind of like with a flow designer but pattern designer doesn't seem to want to do that so some of my steps were not in the order i would have optimally arranged them in but i just didn't want to have to re-create the recreate or have to copy and paste it to a different spot in the order anyway um so this uh parse command output step just invokes my uploaded script and then it parses out my uh it parses out my information i kept it really simple your options with regard to parsing data are within pattern designer kind of rigid so i just kind of kept things really simple because this is a quick prototyping exercise to turn around and provide something to show the customer within you know within a couple days i'm not sure if my debug mode is still running i'll just click the run command on the off chance that it is still connected just to kind of illustrate what comes back so um i chose i just chose the at symbol as a as a fairly viable delimiter just because it didn't seem to show up in process listings uh very much or at all so for my kind of initial poc use case it was easier than like a colon or [Music] i had all kinds of trouble trying to get backslash as a delimiter because that's a reserved character in so many languages so this is what i ended up with and it just kind of pulls each line from the process listing with the container id prepended to the beginning into a table that's now accessible to my pattern and then it saves that data until after it populates the docker container step or the docker container table in step 53 so this is the out of box step that executes and it um i'm going to cancel this because debug mode can sometimes go out and do its own thing for a while so this is the out of the box step that populates the docker container table and so rather than mess with that i just added a step after that which then takes the output of my uploaded script loops through it and looks for anything with the current docker container id prepended and when it finds a match it just adds that to a big string which it's going to store in this really handy column that seems to exist in a lot of our cmdb columns uh cmdb tables it's called attributes i didn't do an exhaustive search to see what it's used for where it's used in the cmdb but i did look at all of my docker containers cis and they all had nothing in that field and so it's an out of the box field it holds up to 64k of data so it seemed like a good um field to kind of co-opt at least for the uh extent of this this quick exercise there's no reason you couldn't create a custom a custom attribute so the reason i did it this way versus trying to push the running processes directly from the pattern is i haven't quite figured out i've gotten as far as finding out that conceptually you should be able to update basically any table from within a pattern what i haven't cracked is exactly how you do that for something that doesn't start with cmd bci um with something that starts with cmdbci in pattern designer you just define a variable that matches the name of the pattern of the table and then at the end it bundles it all all up in a nice ire payload for you and sends it along and your table is magically updated when i do the same thing for a table name that does not begin with cmdbci so the running process table starts with just cmdb it's cmdb running process it populates the table just fine within the pattern designer but it never ends up in the actual table so um again you know short kind of i had a short time frame in which to work i just kind of exercised a workaround so i just sorry somebody have a question no nope that was me oh okay um didn't spill your beer i hope no okay good that's that's the most important thing um so what i did instead was i just took advantage of this empty field that had 64k of storage allocated to it and i just dumped the process listing into that field so that if i go to a container oh no it's just plain old container so if i go to my container listing at the top is my my kind of my test case and you can see it's got a whole bunch of stuff in this attributes column which corresponds to the process listing that was discovered by my custom script there and so that's where the pattern kind of hands things off the pattern populates the cndbci container table and kind of seeds it with process listing information in the attributes column so the way i kind of close the loop and get that information into the table i really wanted to belong in is i use another really cool discovery feature called pre-post-processing pre-post-processing allows you to kind of inject uh inject special behavior before or after the payload is processed by the ire so i created a post script called container processes and i set it to run post sensor and the script is automatically provided a full it's basically provided the full payload such as it's delivered from your pattern and then i just created a loop to kind of loop through it and for any docker container record that's included in the payload i pull out the attributes and if they actually have something if it has something in it then i just kind of go through and parse that out and then generate the running process record for each of the processes that are found on my container and so the end product of that such as is currently if i click into my raw docker container that i've got here is now i've got a running processes tab for my docker container that contains a reasonable facsimile of a unix process listing and so we haven't gotten back together with the customer since we were meeting with them and they said hey can we do this so um i don't really have feedback as to how what what they think of it but i mean it seems pretty similar to what we got in the server side and it wasn't um other than wrestling with some some syntax things which are always uh just uh quite the adventure sometimes with javascript and uh patterns also support groovy under the covers some of the example code that i was looking at was written in groovy so that was that was a a pretty cool adventure as well i haven't really done a lot with groovy and um yeah so hopefully it's it's well received and and kind of allows them to visualize what we call the art of the possible in terms of just showing them hey you know with a couple hours of effort here's something you could extend if you find that you know the out of the box gets you to the one yard line and you need to go that last yard to get to get the touchdown questions okay well that's that concludes the tech deep dive of our segments so um mike is are we just going uh open forum at this point or do we have anything else that we wanted to uh any anything else formal we wanted to cover uh the only last thing really is um uh just the next session um is going to be september 22nd 4 p.m eastern from a topic perspective i think we're actually working on a a log multiplexing solution um around doing log ingestion out of containers so we're working to finalize that right now and we'll get the invites out to make sure that you guys know about it well in advance but um i i think it's going to be pretty interesting uh okay so now let's move it over we'll open it up um for q a and this is the part where ask us whatever you want to ask us if it's related to what we've talked about today awesome if it's hey i want to know about something totally different feel free to ask that as well and i'll i'll stop the recording just so that folks can feel more comfortable if if uh yeah if you don't want your question immortalized on youtube then exactly we got you covered here we go
https://www.youtube.com/watch?v=12aBTxV8ISo