logo

NJP

CSDM deep-dive: availability metric calculation discussion and overview

Import · Jan 30, 2024 · video

I'm going to share with you my presentation this is csdm as we know and when I explain csdm to a lot of our customers I the way I explain it is this is your physical layer this represents what's physically in the data center or in the cloud data center right these are all the thousands plus CI types that you have out there this is more of a logical layer I call it the logical playground the app Services is is logical and it represents the instance of an application it also can mean a number of microservices is in a larger app or it could be also representing a platform with many apps running on it this is sort of that the highest level logical tier that represents all of these components down below how they work together so that's kind of how I present it and in the definition that csdm provides we also explain it that way so just looking at the definition real quick it's a logical representation of a deployed system or application stack yeah makes sense from my perspective too just knowing what how function on the platform yeah now what what becomes interesting and this dichotomy between Technical Services and the Business Services that these I I explain these are typically the piece Parts these are the folks that provide the individual components and we have to use this Dynamics a group as a query mechanism to say well which component which network gear is for development versus production and a query would sort that out and you would associate that with a different offering that has a different commitment the different SLA Ola and so because stuff comes and and leaves the environment uh dynamically in a lot of cases this query keeps the relationship between the offering and the components accurate as they come and go I don't know if that plays in the availability calculation for these types of services but at least from a management capability this this makes managing those items a lot easier it absolutely does especially or particularly I should say when it comes to the commitments you mentioned the slas the OAS we have availability commitments and so you can say for this particular technical service offering we want to maintain this level of availability that's right for these reasons and then when we calculate availability we will will tell you basically whether or not you're you're hitting that commitment so these guys don't know how those components are going to be used they're like I manage 30,000 servers 10,000 are development non-production 20,000 or production High availability or whatever that happens to look like so they would structure these offerings to really better relate how these equipment is supposed to be managed how it's managed and how it would be used in a larger context yeah the next tier up here from a technical service this is more at the the apps and platforms that are are not used directly by the business and a lot of times these are going to be services like ldap these would be a deployment of ldap or these could be managing VMware instances so Venter for example is not used directly by The Key by the business but it still needs to run it still has a set of commitments at the offering level and there it does require at a logical level you to understand all these underlying things and how they're working together to make that thing work so so that's kind of where this comes in and also represents platform I often talk about service now in the context of our platform really doesn't do anything until you use the products on it you start configuring it you start providing access to features or buying new products and layering that on in the structure of a platform the app service represents a running platform platform instance and then there's app services that represent the products running on it we have examples that get into that that's the next layer of Technical Services that I usually explain and you know once you go through this and show some examples people usually understand that I would only say like um this also becomes really important when we talk about the dependencies of the business services and offerings on those technical service offw and the commitments that you're making at the technical level and how those are are you know related to or in congruence with what you're promising to the business ultimately yes and and one of the things I love about what you did in the service Builder is to kind of make this offering front and center from a commitment perspective but I also talk a lot about how we use subscribers we covered that in one of the the previous videos and you can say well there's individual users or a group of users or a whole business unit and if you can relate the groups that to the users individually or the business unit to the users you can then notify everybody at a location for example you can kind of calculate and get yourself a list of all the individual business folks that might be impacted by a service outage there's a nice ability to trace through let's say a specific server or let's say a network uh device that fails and how that traces back to the individuals impact it on of a very mature implementation yeah exactly because maybe the network gear failed or was offline for a period of time but the way that the system was structured or there were things put in place to avoid full impact or or felt impact you know from user perspective and I think that's where availability the nuance and sort of the art of it versus the the direct science is really understanding true availability in the sense of people being able to use it consume it versus something is offline for 30 minutes but nobody was actually impacted it's sort of like did a fall in the woods if nobody was there you know didn't make it sound and to give you some context and really we had two data centers that were identical or close to identical and we load balance traffic between the two so up here there was a network device and we directed traffic to one set versus another set of servers we could make a whole data center out fly and and nobody's the wiser as long as the traffic wasn't high enough to exceed the capacity of of the other data center so yes we designed a app service with the load balancing and redundancies kind of built in we did log shipping of the database to you know logs to be able to keep them in sync and to to queue up the transactions right or the other one was was down that's such a good example yeah so eron um where this becomes complex is because these things are architected differently and and how everything relates to the app service here or how it relates now to the business outage records now I I believe we're we're tracking outage records down here at the individual CI level yeah and you know not only do we track like an outage at the CI level I mean like you can have an outage you know packed D CI directly but you were also able to map an outage to many CIS as well through you know the affected CIS as well so like uh yeah it can have a lot of impact yeah so now that you've got commit sort of over here from a technical perspective right there's the there's these folks that provide commitments on the individual devices there's folks that are you know providing commitments on things like platforms and then you've got commitments to business customers here uh which which you have that I call it the last mile it's like your electrical distribution there's the stuff that connects to your house individually but then there's these big huge high voltage lines that have to that cross the country and nobody connects directly to those right so that's kind of like dealing back here these are the high voltage lines so the question is how do we really calculate this stuff when a failure happens like a piece of network GE which which might affect multiple app services or not and then how that affects the end user and being able to tell them oh you know we met your commitments or we didn't and here's why I see it at least from a history perspective a lot of brms get into these discussions business relationship managers to kind of explain how the portfolio works and and why this this happened or didn't happen the way it was supposed to yeah so from a product point of view what I could do is I could talk to it from a overall use case and and highle product and intention and then Ain will be really good at getting into the nitty-gritty from the product side what we did with availability is in order for our availability engine to calculate yeah yeah the CI but specifically really we're looking at the service offering needs to have an availability commitment created okay so that's a because in the in there under commitments you can put any kind of metric yes so so with the commitment you choose the type and we can show this you know visually in the product but you choose the type and then you set what the the committed availability would be out of 100% okay and so that's your target that's what you're what you're shooting for and you can also choose a schedule so if you for instance want to have a commitment that's for you know your high traffic times or work days and then you have a different schedule or different commitment for holidays or weekends the more we talk to folks though it's everything's becoming so digital and expectations are so high from consumers it's becoming just more and more 24 by S yeah yeah but we do have that flexibility so you could create many commitments that would map to a single service offering and then those commitments can be shared so you can have a series of Standards you know like gold silver bronze there's a a bunch of different ways that you can Define your commitments ultimately though once you map a commitment to your service offering we will automatically generate availability results for that particular offering we do that at different time periods so we'll do it like you know today or last 24 hours you know uh every week we'll do month we do year we do last seven days last 30 days there's a bunch of different reports that you can look at and again we're just doing that automatically and then anytime an outage is generated and it's opened against a service offering you know in this example with a commitment we will take into account that outage time into the availability calculation okay yeah the use case really is and and I will say it's a it's manual right now so we get a lot of questions or we get a lot of requests for an outage to basically be able to roll up from an application service yeah an offering you know just roll roll that up automatically add it to the offering and we've been very cautious to do that and we've kind of push back on that use case because like you just talked about the example that you just talked about an outage at the application service level or possibly you know lower than that in the stack doesn't necessarily mean impact to an end user or that that thing was not available yeah so there does have to be some sort of an assessment done you know like when did it actually start what was the level of impact you know who was actually impacted and for how long um and so kind of unfortunately it's it's it's kind of a a manual process and human assessment at this point and so we we really recommend that what you actually open or create an outage against you think it through basically it's not just kind of a broad brush because that'll have you know a lot of impacts to your results for Avail yeah right I don't know if you work closely with the event management folks but if you're monitoring these individual components I don't know if they're creating outages and that comes into play those have to be evaluated to see if they're really is an impact or not and having a good cmdb helps right having all this structure helps but it like you said I don't know if you can completely automate right yeah not yet and just like you said we really encourage the use of those event monitoring tools and using like you know application monitoring because that'll give you a much much better sense of what's actually happening at the tech layer especially when you want to do like root cause analysis right but when it comes to availability and how we think about it at that higher level it really is more about like availability to the person who needs to use it and so my experience and this is going into my APM days a little bit but when you know you talk to the business about an idea to develop a new application and eventually a new service right that it provides there's a different cost there's a different cost of development and design for high availability versus low availability doesn't matter right I mean there's a lot of groundup redundancy built into the whole thing which costs a lot more it's a lot more complicated but of course you know customers are happier because they would continue to operate without much of a hiccup yeah you know that's such a good point because that investment up front right ultimately hopefully will result in maybe immeasurable savings down the road road and it gets tricky too right when we talk about more of like a degradation versus an a full outage and we really haven't solved for that in our calculation and Aon you could probably talk yeah more about why that is but that's tricky because we do have three different types so you could create an outage that's I think we just call it outage which is more of an unplanned outage then we have the degradation and then we have a planned outage so if somebody creates a planned outage that won't impact the availability commitment because it is something that's planned you know like right Y and you can also create maintenance windows and if something happens within that maintenance window it won't impact the the committed availability so we've done some of that but the degradation is tricky so I don't know eron if you want to talk a little bit about that one yeah I mean so the there's two things that actually came in mind so the first one was like when it comes to availability how we view an outage so for example example you know we have like an outage record that's one thing but when it comes to availability we actually do a lot of um we consolidate like overlapping outages into a single outage right because I think the way that we view it is that it's an outage on that service offering or app service so if you have three outages right three outage records generated over like an hour period but they all overlap you know we view that as like One Outage now what makes that interesting is then let's say if you have like a planned outage right and in a planned outage there was a that's a planed downtime so if a planned outage overlaps an outage and let's say if you're offering right whatever it is is currently having an outage two hour long outage but in that period of time you plan to do like a planned outage of some sort that would actually split the outage time into two outage periods in a way right like if it ran right in the middle because it's one was planned but but then it's really you had an outage on the beginning a planned outage and then an outage at the the end and so in terms of availability this is like where a lot of this complication kind of nuance comes in you know we actually view that as two outages right yeah two outage periods two times where outages were down um and that's some of the Nuance with planned outage but when you get into like degreg yeah that's that's hard because what is that right exact like yes let me talk through some scenarios that I was accustomed to in my history okay and I'm going to go into some csdm examples so I'm going to go into a microservice example because I think this easily conveys from my own experience kind of what what can happen so actually let's go to this one first this is a decomposed Mal so this is just a great simplification of something I I used to have to architect and build and support and in this case I'm going to start with there's three key modules here there's a tax calculation module there's a currency conversion module and there's a customer management the main module of this particular application that that manages the customer interaction and and in this case it's all part of one business app it's called online sales management and there's one app service that stands for the whole thing okay now under that each of these products or microservices when you deploy them independently and they're all interacting with one another so so these things just kind of interact and and I don't know if you're familiar with microservices architecture Aon but I've worked with this architecture since the early 2000s and you know this this is just the way we're doing things way back then I didn't we didn't call it microservices but this is this is how you had to do it it's funny you wanted to scale certain sections differently than other sections and then of course some of these things can be running on shared infrastructure so one piece of infrastructure goes down it might in fact multiple microservices in this case for online order management you might have the main cust customer management module up and running but you can't convert currency or you can't calculate tax to finish a sale but you can have let's say the currency conversion fail but you can still calculate tax so this would be like a partial outage scenario where you could still process business but you had some cleanup or you had some some certain aspects of the business that didn't work right so the the the business service over here online ordering was still up and running right but a portion of the the functionality was suffering because there was an outage of the currency conversion this is a very realistic scenario my little microservice over here was the the you know apply for credit get credit buy something sort of module and then there's also the the trickiness or if if maybe these are all available or they're they're quote working but one of them's really slow and so yeah it times out or like one in three transactions fails like those kinds of things that are really hard to actually capture like Aaron said what does it actually mean is it slow is it not fully completing and partially completing that yeah that's like how do you yeah exactly with the degreg like how do you categorize those things into a calculation that has definitely been the trouble with that well I have to say this this was another use case that I had where we had this loans service which I my team owned and you know this is what happened it start the it started to queue up so all the traffic would queue up and if we were if we took more than 5 Seconds to respond the the transaction would would not work you know what was wrong there was nothing physically wrong with anything it was the backup Services were running at three in the afternoon which is where the traffic was ramping up so you hit a certain threshold and it was just C it was just yeah the thing would slowed down and we would lose transactions but yeah I mean it's a real I mean it was a war room nightmare 3:00 recurring nightmare yeah what we've tried to do to to help not solve for it but provide some kind of of a a tool to track it right is is give that type the degradation type so at least can be logged and it can then be mapped to to all the different you know CIS in particular that that are being impacted and then hopefully can help inform analysis especially if it's Rec cing yeah to try to do that deeper dive and so at least you have a record of that but we haven't been able to kind of crack the code on how it should be part of our calculator so to speak yeah it yeah well I wanted to walk through this scenario because assuming you have a microservices architecture like this and you you follow some naming conventions to describe what these things do and you could put a number maybe a number on what's the value of this thing if it goes down completely it's oh we can't do XYZ we can't convert currency or we can't provide a loan to customers or whatever it happens to me so you can now say something about online ordering was degraded because we couldn't calculate currency conversions from three to five or whatever the heck was going on yeah no it makes sense it's interesting for sure um and I'm just trying to think with like you know like I said when the way that with availability there's basically two entry points right it's the commitment Associated to you know an application service or an offering but then it's also just you know the outage that itself is generated and I'm just thinking with an outage like I kind of mentioned before you're able to associate with multiple items but then also having that further relationship maybe of like like you said how these things are related to each other as well and impact each others is interesting Yeah It's tricky in that one example you gave right Mark it's yeah well because this we had the failover the redundancy it didn't actually ever impact the end user and so we're not counting that as it being unavailable right but yeah right it's a good point um yeah but you can I you know that's one of the things where you can make that call based on the situation here right and then measuring actual user complaint that anybody know these transactions all finished right I mean whatever was going on still continued to to work as an architect you try to build in these redundancies and fail safes along the way but you don't get everything I mean nobody predicted this 5-second thing was happening right it was like how do how do you know you know they just kind of gave up saying thinking we're down but we weren't down if they waited 10 seconds we might have you know continue to work the tuning the architecture operationally is is a bit of a bear to be honest because you have to have visibility and what's going on I I I think it's a real world situation I'd like yeah I think it's a good way to kind of explain and explore the different challenges or sort of the Nuance yeah when it comes to calculating these things and understanding impacts something is available Andor impacting and and the end user I mean ultimately what people want to know is you know are are we meeting our commitments here is sales still working and if it's not why what happened was the cost of that so that you can mitigate the worst case scenario is that this thing is architected to not be redundant at all and if it goes down the whole thingap is kaput right so you got to look at your weakest links and then find those and and act on them right exactly and the data and if you're doing this calculations that data I think also can help because if that availability can inform and maybe also help make a case to take more time and fully Arch ecture system with some of those redundancies right if you can actually show hey the way that we've designed this in the first place is costing us this amount of money and we can tell you that because of this availability calculation in you know accordance with also the cost of it not being available and then actually make improvements you know I think that's the whole idea right we want to provide data and we want to provide visibility into that for continual Improvement and enhancement and a bunch of different ways I would say the better granularity you have on that data you know what happened and that's what I was asking about monitoring and sort of that itom visibility or itom health kind of tie in to and then when it comes to like the consumer though that's like when if it actually impacts them or not I guess that's more where we see our availability come in our availability calculation specifically yeah I was I was thinking maybe um we could you know kind of briefly show the availability product itself and kind of talk a little bit about how it's calculated and we can y do it at a higher level all right so um I guess I'll drive and then Kaitlyn you can tell me um can I guess chime in since I'm sort of have this on my screen so first thing looking at here I just want to since we're kind of going from this high level kind of want to start there so you know we have our in this like demo instance I have you know we have our sample portfolio and this is just a service portfolio you know my world this is kind of where everything begins um for least like things I work on you know we dig down to the T Amino print scan effects we have the service printing services and then we where the first step to this availability process begins or where associating the availability begins so get to this first offering on printing sorry all right cool so so when looking at like an offering you know we have a service commitment Tab and these are some like I said this de so there's already some service commitments associated with this different types I'm going to actually go in and create a new one sorry oh um okay go popped up so in our case here you know we'll just do I'm going to name it some really basic but 99% availability this priority exists on it's demo data but just create one for us um on top of that I think Caitlyn you mentioned this earlier about the different types of commitments that you can associate and so you know we have um all sorts of different ones and I think on this instance I have it's pretty there's not a lot of plugins installed so you know another common one you'd often see would be like SLA let's not get into that but in our case here we're going to bring in availability you know you and this is kind of what K saying you can describe the percentages um Precision um there's the concept of SK schedules and so these are kind of like schedule tables so these can be anything from you know 24/7 weekdays America I mean there's a there's a lot of options here we have um and a lot of the rest of the fields here are really just for kind of noting they don't actually um do anything than time zone does do something but um like breach penalty this is more just to track that information it doesn't actually go into calculation okay so so interesting you could say that the availability here would be like for high like like my my online store example right you measure your downtime in in millions of dollars per minute you could put that kind of detail in this totally it's not used to it's not used like you said to calculate but you can use it to create a report to kind of say oh that was a a$ five million doll you know outage or whatever yeah yeah exactly you can use it kind of like note those differences and kind of deal with what you want but what does those matters like the percentage Precision Time Zone things like that do play an impact in the availability results so actually do you want to talk a little bit about how the Precision works yeah so the Precision is just like the it's it's just for determining the percentage availability so if I did 99.99 and I actually care about that like the 99 the 0 99 then I can decide the number of decimal places that matter or so that's the decimal place sort of precision I got it yep that's it it plays a role in the calcul the basically determining if the availability like the outage amount has surpassed the percentage if that makes sense got it cool so I'm availability and I'm call store that's our example I'm just gonna do that so it's easier to track um all right cool so that's created let me go back all right so now we have this one and there's a couple things that kind of happen in the background here so one thing we do um is we do generate availability like historical availability we actually have an action for this too if you want to calculate it for like the past year so if you did have like existing outages but you didn't have availability records yet you can do you can calculate the past that is an option oh wow would that be helpful if you're trying to finetune some of these calculations right recalculate F tune it recalculate type of thing exact yeah so if you're like adopting availability like you're adopting um like if you have existing history of outages but you don't necessarily have like them Associated offerings because you're not using it in that way um do that and also it's useful for recalculate in case so there's kind of I guess a debate here right of like um you know we report on the events that are occurring at the time right or whatever the outages at the time do you want to update those if in case there was errors right and so yeah having the most calculations you know you need to rep for example calculations you could do that too by recalculating basically but there's even more to that I'll get to a second so okay when you go to availability search for you're one thing this platform specific there's availability module and then there's an SPM availability this is important because in SPM availability we are mostly talking about associating availability with offerings and service offerings specifically yeah so I just wanted to jump in real quick and say that with availability it started in service portfolio management module with outages and the you know creating commitments against the offerings and we were receiving requests more and more frequently to have that functionality for application services or like other CIS and so we made that available U but that's why there are two places you can find availability yeah exactly it was something originally for SPM that we extended out and so it yeah it's basically been kind of extended to work with application Services there is another way there is a way to make it work for other CIS but we ship about a box for application Services now I believe was as of let me actually some reach out asking about it okay let me actually I think Tokyo I think well yeah Tokyo okay so to kind of show so this is now this is what availability looks like right like and like fundamentally like the records that we generate the core like this can obviously turn into into reports and be monitored with indicators and kpis like things we do in DPM for example but um you know just availability itself we B basically generate these individual records and Records basically are made up of a start date end date associated with a type and so that's important because we have you know like Daily records we have the annual availability monthly and so yeah sorry so besides this happens you can also see we have like you know total availability percentage downtime yeah so we can kind of see the breakdown here so these are all the different types of Records your your annual records your daily last 12 months so we have like rolling so any of these last ones are rolling values that are updated daily and then we have our monthly oh so you know um the days of the month and then the weeklys as well and so these will get generated nightly so now this is where kind of avability can maybe start I would say getting a little confusing is when are these records generated so the first one is and we kind of saw here when you associate a commitment to an offering for the first time you get an will have kind of like a historic generation for the last year or so we generate throughout all the records that would have happened so that's one form of generation okay the second time is we have a nightly job a scheduled job but there's basically a calculate availability schedule job that runs at midnight down out a box and that will calculate yesterday's data right so if it ran so the one that ran today January 3rd at 1:00 a.m. or so with actually be calculating for January 2's data and if it was the end of the week or end of the month or end of the year it would also generate the weekly monthly data now there's an aners to this too now because with availability up until when we're more recent releases we would only calculate when the month would end or like the year would end and so with the availability calculator V2 which is a a brand new for any sort of new customer they'll have it but for existing customer if they want if they would prefer to have the current month constantly being calculated current week the current year being calculated instead of the end of the month week or year then they could use availability calculator V2 it's documented on the on our docs and how to turn this on it's basically just a system property flip but that is kind of like the a lot of the feedback we got was that they prefer that customers prefer to actually have the ongoing monthly calculated um but in the previous ver version you know we kind of believe that you know you should only calculate it once the period's over but people actually like having the on goinging got it so then so sorry so availability calculations nightly you'll see these records get generated as well as and the kind of the final case and I'll show it here and in this case I guess in this use case you would say this outage you know we have this filled personal printing or configuration item and we associate with the personal printing directly you know if you were to register outage in that way I'll just say typically people will just like start their out maybe I should do that but you know I'm not going to do that actually I'm to make this for easier to understand we'll say I was logging outage that actually happened you know yesterday at 1057 to 1257 and I'll show you why but obviously most cases here would probably be you would create the outage and a period of time would go on and then you would put the end date and then hit submit so now when you're look at the outage records um we have our beginning end date duration um and then we have this affected CIS the affected CIS would be a way you can associate other CIS that would also be impacted by that outage and that would actually reflect on the availability results for the existing um the configuration item also too now when you say effed eyes does that include app Services yeah so you can actually kind of throw anything in this L list but application services yeah are one of the ones that you can add but Al specifically will be the ones that could get their availability recalculated or calculated if they had the commitment Associated to them got it app services and service offerings as long as they have a commitment we will do the calculation and you could of course put anything in there in terms of what's being affected just the availability won't be calculated if it doesn't have a commitment okay so I'm going to kind of quickly go through this so now this is a result record and so this is a result record that did have an outage that outage we created and you can kind of see that in this outage is during interval so we can see that there was an outage here from you know that one I generated one that we created but we can also see now um so let's start with total so the total values here are actually the total availability this doesn't include the schedule right as well as the percentages so here we're basically just seeing like over that day or in this case it's a month right because this is a monthly record and this started this is for the month of January 2024 so this current existing month we can see that has a total availability percentage of 9.73 there was one outage now when we can take in consideration of the commitment which you can see this percentage is a little bit lower and that was because for the service commitment we did 8 to five weekdays so the percentage is going to be a little bit lower not as not super noticeable because this is not including Saturdays and Sundays right inside of the availability percentage the total is less than 100% or 100 or whatever the the number of hours there is in a month yeah and then we're also being able to see to that commitment how much downtime has there been so far this month so two hours six minutes yeah and then once again we see the total the periods and this is important because when we calculate outage t uh periods it's that includes overlapping outages and the complications that come with that and yeah so here is kind of like that end result and this you once again if you associate an application service they would also have their own monthly records too every single one of them for every single availability commitment it's all broken down individually got it for the customers that are using this are they really getting into these numbers at this level and and and how are they using it just very curious what your feedback has been what your experience this is probably one of our most used maybe is the right word or highest interest uh with most of our customers around DPM and and service portfolio management for for a number of reasons but availability is really important for for most customers and so we get a ton of cases we get a lot of questions and it's probably also one of our most confusing features so the customers are really getting into the details we have a lot of like sessions around between periods you know if an outage occurs in one period of time and it overlaps in a second period of time or challenges around time zones and calculating availability in various time zones there's a lot of nuance and there's a lot of specific customer requirements you know there a lot of them are doing it manually or using a different tool and so moving over into service now and getting that visibility and getting this working for them and making sure that they're getting the right numbers you know that's matching how they're calculating in other places so I would say with a fine tooth comb yeah yeah I could I could definitely see that there's a lot of moving Parts here you got to get them all right yeah and and from like kind of a support standpoint yeah no the and the people who do use this the customers that do use this I mean they they rely on understanding the numbers and the numbers being very accurate so that's a very important part for us is to one understand why this commitment Avail percentage is less than total availability percentage because right it's 99% it should be the same but making those details clean and or clear and the differences between all these values um you know because it sometimes things don't look like they line up but um there's you know usually that nuance and explanation that can can help clear that up for customers and users right and then of course you know tracing that to actually the the data that's in the details and maybe explaining it you know to somebody yeah I'm I'm curious why is this so important for those customers are they are they being held accountable is this penalty real for them yeah so I'll say either yes like working with vendors or working with other teams um or also being able to report back to the business so just like you talked about being able to to quantify and and more and more businesses kind of require requiring that transparency and and that reporting so many customers that I talked to when they start with like digital portfolio management for example it's they start with their service portfolio and the availability got it um maybe you could just jump into DPM real quick and just show so one of the things that we've done is you know there like aerin said you can create a bunch of different reports using this data to get exactly what you need but we do provide some basic outof boox reporting in digital portfolio management for the service portfolios uh kind of showing how availability is rolled up at each level of the like from the offering to the service from the service to the taxonomy node all the way up to the top so it's a really great way for customers to like drill down through the portfolio to see where uh there may be like dips in availability in particular you aggregate these these numbers at the portfolio level right so yeah and then be able to like drill all the way to the outage record you know yeah be able to see the trends and it's like kind of progressive disclosure so we do provide that out of box like I said uh but then there's a lot of different things that customers can do you know with with reports on that data to get exactly what they're looking for and one thing that we did do is Aaron's showing the different types of portfolios so we support application business application portfolios and we will show availability calculations for the business apps and it's based on the application services and so if you are tracking availability at the app service level and you've commitments defined there uh we will roll that up to the business application Level um which was yeah that that's interesting you know it goes back to this whole Synergy between the service and the application definition they're slowly becoming one right there's yeah we're getting there okay well you can see the service availability you there we go and so this is interesting right because this is a taxonomy node so we're actually seeing availability records for all of the service offerings that make up this essentially aggregated score wow it's a lot of them too yeah so these and these are calculated like you said daily or monthly or whatever you've got them set up to do so these are specifically Daily records and so obviously this is 100% we no outages on here you can imagine this if this wasn't a a successful taxonomy node that had zero outages so realistic um that would be you would see the you see be able to see this history right and you know what's really cool with like DPM is you're able to really set your like you as well so like you if you really care about availability you really want to be able to dig through it and you really care about like that cmdb yeah you can you have that option inside a DPM got it yes well there you got one for 99% one quick Nuance too so with this needs attention because we do get questions of like the difference between the needs attention and then what you're seeing in these in this performance snapshot okay so like Erin mentioned the availability the one that we share I should say it runs you know every night so you're getting yesterday's availability which wouldn't include for instance an outage that was created today so the need attention is is intended to give you those records that are happening today or now so that you can have that visibility okay yeah understanding the differences um and but also understand that kpis themselves are also kind of like an aggregate over a period of time right like you know in the need attention we only have one critical incident but in open incidents it's there's 420 and you can dig into like you know what this actually is we also you showed the commitments um really just on the offering views so you can see the commitments that have been set and also whether they're being met um kind of each each month so for both SLA commitments and availability so it's an additional View and DPM oh okay I'm familiar also with service Builder is this detail now available in service Builder or can you launch into availability from service Builder what you can do in service Builders you can map the commitments once they're defined yes correct all right so so so somebody should know what kind of commitments need to be done over you know over everything and then service Builder is where you can select the ones that make sense for each service exactly kind of the sequence yeah okay okay yeah yeah it could be a different Persona as well so yeah there's a lot of reasons why um but exactly you would Define the cment outside got it but once Define then the service folks can define a service using the appropriate commitments yeah perfect that makes a lot of sense to me uh so there there's a lot of I would say setup that customers would have to go through to get service Builder to really sing for them in that regards correct yeah it's the idea with the service Builder it's it's for yes yeah like more of that end user to kind of pull the pieces together yeah okay yeah but and really what it comes down to yeah you just need to find your availability commitment you associate your availability commitment to an offering and well well the history will generate but then you just kind of watch it and the next day there'll be a new record and the day after that there'll be another record and V2 just remind me again it's already there they just have to go to a system property to turn it on is is that Al also enabled by default for brand new installs is it already is it defaulted V2 uh yeah exactly it's it's default now for new new instances but for existing ones they would need to go to this Comm SMC availability B2 we can access this property admin can and then you just kind of flip the switch to true yep yeah and then the new way of calculating is implemented on the back end they don't see any difference in the user experience the yeah pretty much the exactly it's all done the back the big difference they'll notice right away is we moved from a pattern there's two patterns we moved from so the first one was we would calculate the previous month when the month is over like the current month would only calculated when the month ended we switched that so now we do kind of like an ongoing calcul or like the current period records are being calculated the other big change in the old calculator we had this concept of ongoing outages and the idea would be if you had an outage that begun with no end date if it flipped over to January 4th we would show that there was an ongoing outage and customers didn't like it it actually really confusing it was definitely some issues around like expectations so mostly we go people were like eh we would prefer it if once you end the outage then you just update the calculations and so that's like the so we basically switch to that which is that you know there the an outage is not calculated an availability record until you put that end date in until that end date is there and to build on yeah to build on what you were saying Aon just so it's it's more obvious if an outage is created at like 3 pm today in the version one calculator we would recalculate the availability for you know today but we would assume that the outage would be open through the rest of the day because we couldn't guess when it would be closed so then we have to say well then the outage is going to take from 3: to midnight basically which is not looking good on that day's calculation no no and it's misleading right because it's like well that's in actually oh and then by the way because we also do weekly monthly yearly and like last 30 days last seven days all of those would assume again that the outage is open for the entire period so to the end of the year to the end of the month whatever it is so instead now then once it is closed we would recalculate and everything would be updated to reflect the new time but for a while it's very misleading and it's a little offputting you know and so then with this version two we don't make that assumption but because we can't assume that it's only going to be open an hour or two hours we just don't include it in the calculation until there's an end time wow okay so there there's a lot to it but now I think you're at that point where you're now sounds like you got a lot of customers you're getting a lot of good feedback to you're already on ver version two original availability I mean it's it's been core and it's been there for I mean I the oldest code I've seen is over a decade so it was yeah it was due for a little bit of uh espe with overhaul yeah yeah all the feedback we we get with it that it was definitely a good work doing I think I hit all the big points uh there's so much more I could show I could show like planned outages but that's a lot I feel like we already yeah yeah these things are more you know high level scratch of the surface and getting folks pointed on the right path on how this comes into play right with regards to the the whole model we want to engineer these systems with these numbers and penalties at the beginning of the process so that we properly engineer them right so I wanted to thank you to for meeting me and kind going over this csdm and sort of how this availability calculation so a little bit about how digital portfolio management uses this data to see it at a higher level uh this is very valuable I know our customers would definitely definitely want to take advantage of it if they knew a little bit more about how it works in the context of the bigger picture yeah thanks for the opportunity it's always always fun to chat yeah this is think first half was very interesting for me too absolutely let's do it all right take care bye

View original source

https://www.youtube.com/watch?v=GizZ36ysEsA