Health Log Analytics (HLA) Expert Panel - Strategy & Implementation
Einar & Partners
·
May 04, 2023
·
video
[Music] thank you well um hello everyone I um warm welcome to the very first panel of I9 Partners um demystifying health log analytics and predictive AI Ops um so first uh really ever expert panel of us organizing this and we are thrilled to have you here today and over the next hour we will discuss this exciting topic on leveraging AI and analytics for log analytics and how to transform it operations with it as it can have an immense impact on how you manage your I.T operations and so let's get started with it um for people who don't know me my name is akif basser and I'm the RND lead at I9 Partners Research Unit I don't delete on AI Ops and data science um I'm your host today I will moderate panel discussions and very briefly about me I have keen interest and passion in Ai and data science and I love to create industry-wide benchmarks for item and AI of solutions of course I'm not alone today I have two great co-hosts with me who have extensive experience in implementing and even designing blog analytics Solutions uh so Alexander ARCA and Ash Foxon with me so could you please introduce yourselves of course Ashley you want to start or should I now like you're first on the list so yeah all right all right thank you um Aki thank you a lot uh thanks very much for having me here um yeah my name is Salix um yeah what I'm doing we founded the company a few years ago helping um our customers to establish actually lock analytics um we mainly focus on customers that do not have any kind of lock management or lock analytics yet um and basically guiding them through the Journey of setting it up how to set it up what are common pitfalls um to to be aware of and how to end up with a log analytic solution that is actually helping the specific business of our customer at the end and yeah a little bit background about me um I did my PhD in this uh in this area where I researched basically um different machine learning and AI methods on top of telemetry data of software systems and one Telemetry data type obviously as you know it's log date and therefore yeah that that was um what caught my interest and we pers or I tend to stick with it now nice nice and great to have you thanks and uh Ash who are you yeah okay I asked myself that question every day um so Ashley Parkson uh and the domain architect uh for digital operations at National Grid we are so my I've been with National Grid um six years coming up to and for people that don't know what National Grid do we are a UK and a US business and uh we're in the utilities industry we generally provide gas and electricity either through transmission or in the US through uh retail as well um we digital operations are a function within uh global technology operations where we we have a responsibility across cmdb Discovery through um to observability through to Ai and also finally automation it's a new function with inside National Grid um we've been running for about two years now um so we've literally had to design it from the bottom up and observability is probably one of the observer observability in AI is probably one of the big things because it's quite new to a lot of people in National Grid and it's definitely new to a lot of the Legacy applications and infrastructure that we've got knocking around so it means a lot of things to a lot of different people and uh yeah it's interesting you know um how systems are built without observability and how they've survived in a digital world without observability or or log analytics being being around I don't know how some of these systems have carried on working but yeah we're gonna fix that now guys yeah very interesting and I would love to we will hear that more about uh um later in the session um we have a lot of ground to cover um let's set the stage for today's panel discussion and well what you can expect to learn um doing this discussion um like like what you just mentioned that you journey and your experience with log analytics for the health of your system I.T systems um we will touch upon that and and what value it brings to it organizations so we will first take such the panel discussion with that you um log analytics Journey uh also your journey Alexander how um um what value it brings from you perspective to it organizations it's always too good to touch upon that aspect when you consider to implement log Analytics and of course we will touch upon what the capabilities of log analytics are and also on how to act on these predictions like the idea of predict having predictive intelligence how it impacts the it organizations operating model and and also to further for those who are more interested in local analytics and how to optimize their experience with it we will also touch upon how to select um or what data sources to select to ingest in the log Analytics very briefly uh um that being said um I also want to note that this panel discussions will be recorded and published later on our YouTube channel and further Set uh expectations for the uh the audience um it will be mainly a deep dive sessions a kind of a guide on the uh like I mentioned earlier on the AI Ops HLA Journey it will be mainly like you already talked that observability so it will be mainly on the observability side of things and it's it will it won't be focused the discussions won't be focused on technical configurations or complex formulas behind HLA and also the discussions are targeted for people who have never worked before with log Analytics as well as to people who have already implemented or used in log analytics and once and are curious to optimize their experience with it and there's one thing though um I want to mention we are assuming that the audience know the very basic concepts of aiops or and observability like um what log data are and how to uh how it's being analyzed uh but with that being said before I give the stage to you Alex and um as I want to um mention um that the essence of log analytics um correct me if I'm Ron Alexander it's making sense of complexity and Madness and uh uh also the discussions I had before with you Ash you can you have large volumes of log data Millions from various data sources and also the different monitoring tools and we try to make sense of that with a set of algorithms of AI machine learning algorithms in an analytics engine and then we hope to Target and or have an outcome um where we can reduce the downtime have a faster root cause analysis reduce mean time to recover and Noise um but more importantly also to prevent incidents and critical outages that's uh also very important and I think you will tell more about that what your experience are ash and Alex and at the end have also deeper Insight in your I.T systems and applications and services of course that all being said um it's finding this magic and the Brave New World of AI let's say um with that being said I want to give this stage to um what's your journey and what your experiences are ash and Alexander maybe starting with Ash with you yeah sure um so so almost a year ago just over a year ago we had a customer a customer of ours in internally at National Grid come to us with a complex problem um they were the product team for a a website an external facing website which serves millions of customers in the US and their complex problem was that I was distributed system uh they had between 8 and 14 different log sources so if there was an issue uh with let's just say payment or billing inquiry or account lookup they wouldn't always know where that problem was or what the root cause could potentially be and given its dispute system we'd have to stand up a war room of eight to ten developers or Engineers to go and investigate all of their blogs for their respective systems it's all in a little black box siled data repository might be sitting on the survey might be in a repository that they have access to but nobody else does uh the pro T Masters can we aggregate these logs into central location and search them so it's an engineer from let's just say the API team can search the logs across all the systems easily and find correlation and causation by aggregating the logs that led us down the path of using servicenow health log Analytics and there was a bit of infrastructure that need to be stood up first so as you know you need to deploy servicenow mid servers to ingest the logs and we effectively took the eight data sources ingested them into servicenow HLA um and I think we're streaming something like 12 million log lines of data every day into servicenow HLA and what that is has meant for the the product team is that they now get alerts or incidents based on anomalies that the product is detecting if we look at what happened before that um and on the slide previously we talked about reducing the noise them Engineers were getting thousands of emails a day to sift through and if you're getting thousands of emails a day what do you generally do you ignore them I will wait until somebody screams and says there's a problem so you automatically remove that proactive monitoring because who wants to look at you know you've got your normal emails as well as these alerts coming through so out of the thousands of emails Engineers used to receive every day we we now have um alert mechanisms in in servicenow and we're down to tens of alerts that are significant so we reduce the noise and we're now just looking at significant alerts that happen in in that system what that's what that's really meant for us is um not only are we getting alerts but we can also look at Trends and Analysis so we know if there's a signal going off in one system that could have an impact further up the chain or down the chain which could cause an incident in an outage by making them alerts become surfaced to the right people at the right time and we're not doing these War rooms we're able to react to early signals to prevent incidents from occurring that's been evident because even over the last six months the team have been able to prevent priority incidents from occurring by looking at these alerts that have been generated that effectively is helping our customers have better trust and um Natural Bridge reputation is maintained people can go and pay the bills a lot easier there's no downtime the engineers and developers can now Focus that time on innovating the product rather than trying to firefight and fix things that have gone wrong elsewhere yeah if we talk about the journey I'm not sure you want to touch on this later but and Alec Alexander Mike attests to this as well one of the biggest challenges I found our developers Engineers faced was the logs weren't always constructed in a way that made health log analytics uh applicable there was a lot of um rewriting of Vlogs changing the logs inserting correlation IDs there was even some spelling mistakes in some of the logs where correlation was spelled with just one or rather than two yet the other system added with two odds there's been a lot of redesigning how the logs are implemented and constructed so HLA is not a silver bullet that's going to fix all of your logging problems so let's just be clear about that to start with because it's not you've still got to do the groundwork first you've got to make sure the logs are in a readable format uh that they can be ingested and also log shipping is is a big factor because you might have a legacy system that hasn't got a log shipping method so it might be sitting in I don't know um the route of a folder on the E drive or the D drive well how do you get that log into health log analytics you can deploy agents obviously like um the agent client collector for logging from servicenow or you can use other methods um but it's you know you have to plan that and architect it correctly to make sure a is shipping the logs B the amount of logs that you are shipping are sized correctly for the platform you want to ingest them into and then secondly all you get is the data that you're Gathering actually usable I don't know if Alex wants to add anything onto that but but it's you know it's not a silver bullet you've still got to make sure your engine isn't developed and what we do we actually ended up writing guard rails and principles so that if they're bringing up new products the developers Engineers know the guardrails they need to work in they know the log formats that they need to adhere to they know they need to be shipped um to a central repository and that you know that could go to for example go to Azure log analytics first and then from there it can go to servicenow or it could go to Splunk it could go to wherever you want to as long as it's accessible via your AI or l-flog analytics platform then you should be okay but yes there's a lot of lessons learned in there yeah great uh we will uh delve more into that later in the discussions one of the biggest things as well I guess is you know the user adoption I think we'll get on to that a bit later as well we're getting teams to change their operational behaviors and practices to now not sit there looking at folders full of logs so just wait for the alerts and instance you know it's a big cultural change that isn't always to explain how you do it I think so too and actually kind of trying to predict uh whether there's a problem issue that's the magic kind of uh tough mentions earlier it's um a great uh great insights and thank you for sharing that um so how about you Alex what are your thoughts yeah yeah I mean there's I mean the insights were very interesting yes and I think you um I'll probably uh repeat a few things that you said already because they're um and they're on point that's exactly it's also um it also fits I would say our our experiences that we had with our with our customers it's very similar a lot of similarities there but um yeah maybe before we start um to or before I start to explain the log analytics Journey um a few words of like caution maybe because I'm um from the customers that we work with uh you can expect that they are rather startups basically we focus on series a series basically companies that are just establishing lock management and log analytics in their systems because the systems probably grew um so large that they are not able to handle them anymore by manual analysis or manual review of locks and therefore they're now decided to switch to a more um more uh structured way of analyzing blocks and establish observability in their system therefore probably the things that I will say are a little bit biased towards that but as as I said there are definitely definitely intersections between what you said Ash and things that we learned so far maybe can you switch the slide then yeah just just a few things that I learned When approaching uh these these companies that are never did anything with locks or structured log analysis beforehand but usually they saw that there's a lot of potential with lock analytics and the benefits they sound really good really nice and they would like to have it right but then of course from technical perspective and from process perspective when you want to set it up correctly right you need to do a lot of homework you need to cover a lot of things prepare a lot of things before you actually can establish log analytics and this is what I summarized here with lock management and whenever we start to explain what's needed in order to have block analytics then people are not that um not that euphorical anymore as previously um so let's switch again okay yes oh yeah and therefore I try to analyze or I try to from my experience to digest or distill a little bit three major challenges that we face when when talking with customers when they want to establish log analytics um uh next slide please yeah and uh the first one is basically the convincing right so their technical team uh and they would like to establish lock analytics because they are not able to handle the complexity of the current system anymore by hand and therefore they need some structured way to do it but of course they need to convince uh the the management layer right the business layer and therefore the convincing argument there is always usually return of investment if I as a company would establish no lock analytics log management what's the expected return of investment for that in one two years whenever and this is very difficult to estimate actually for log management and there I try to um to propose it to four examples um where we see that return on investment can be calculated but it's still difficult um the most or the left part is security compliance um these are this is the easiest way to do it right because whenever there are certain regulations so the company must follow Also regarding log analytics and log management then there's no clear argument for return of investment because if it's an official regulation then you either comply or you're out of business and this is um this is the easiest part for us this is usually out of scope because as I said we usually work with younger smaller companies and there are not that many um uh that are not that many compliances that they need to fulfill but usually for them the other three are more important the second part is the Improvement of productivity right so whenever you establish log analytics you will probably or the team your operations team or site reliability engineering team will spend less time on troubleshooting you will be able to do the work or whenever problem secure you will be able to solve them faster however just establishing log analytics and log management for example with the platform us mentioned servicenow there are several others there as well also open source solution it probably won't improve the productivity of your team immediately you will probably need to First teach the people and to onboard the teams on how to how to actually use the new tool sets they need to be integrated in the daily workflows and this takes time and yes Ash said perfectly it's a cultural change you will need to um to use this tools or the teams to use these tools in order to get the benefits out of it and therefore um even if this succeeds right even if you are able to improve then with regard to the return of investment you will need to do some some decisions you will either need smaller team sizes in order to manage your infrastructure which will save you costs or you can re-utilize the demand power that is now available to do something new or something um more uh you know more Innovations in your company um in there again it's um for every business and for every customer the calculation of the return of investment in the area is different so there's no no common formula that you can apply in order to calculate it um the other two parts the Improvement of slis I put here also slis probably everybody is aware of probably of slas Rights service level agreements and with our customers we are usually working with slis because whenever your system faces and customers like a web service or some some some app that you're providing to your customers then you're not agreeing on certain slas but your the health of your service will impact the user in a certain way right and slas in this area are indicators that are indicating the health of the service that the consumer the customer of your services actually care about not the inner workings which is very important for us to to distinguish and there when we improve this kind of slis of course we can reduce for example the downtime of our service or the improve the mean time to repair whenever a service level indicators are degraded but here if we want to calculate the return of investment we would need to know the cost that for example down times produce or bigger companies probably there are such calculations but the customers that we are working with usually they don't have this kind of insights they don't know what will happen when their service is not available for one or two hours so they don't exactly know the cost of that and therefore the calculation then uh when reducing these kind of times is difficult and the last part which I personally like most is the release was confidence but whenever you develop your your service when our customers are improving their products improving their apps and web services they um still I would say they still struggle with releases so whenever something needs to be released it's difficult because it might break something it might cause impact on your slis and with lock analytics it's possible to verify that your new version or the new feature that you introduced in your in your software will actually behave as it should in production so you can do smaller and faster releases and reduce the time to Market and here their calculation of Return of investment is is easier because usually it's very important for our customers to be the first on the market and to push features as fast as possible to the market and this will or lock analytics usually enables that in needs and the downtime cost was estimated by a gardener and like 100 000 exactly smaller an hour like um can be very expensive yeah yeah you're right you're right however um for uh these are more or less averages right so whenever we approach yeah whenever we approach a customer then uh we need to to take a look at the specifics of this customer and then calculate it there and the Second Challenge which I would briefly mention is that the lock management itself has a very very bad uh total cost of ownership scaling so here it's a very rough formula that I basically took from Ben single man his former Google employee and now has a startup on on traces unlock analytics and here the calculations uh for locks is approximately the transaction rate so whenever your software is doing something rather than it's doing something when there are user requests fired against your application or against your web service and this is basically the transaction right so when your amount of user grows in your system then the transaction 8 rate grows as well then the number of services you will probably push new features to your service and your system will grow and this is where the number of services will grow then you have the costs of the network and the storage of the log messages and the retention period how long the logs should be should be actually stored and this costs a lot of money and whenever you for example you want your system to grow right you want more users you want additional features um but on the other hand it will cost additional costs with your log management and therefore it's very um it's very important to know this beforehand and to prepare accordingly I will mention a few specific things in the next slide um like if you switch and this is basically um where I would highlight the complexity of log management and log analytics all the steps that needs to be done in order to come up or to end up with a good log analysis and it starts right at the beginning right that the left part actually the source code so we're working usually with companies that are developing Services themselves and they have in-house developer teams and this is where the journey starts so a developer needs to write good lock statements in the code in order for the logs to be analyzed later in the log analytics system and Ash mentioned it as well right here we need to be careful when we do when we select the log format we need to unify it across our teams in the company we have the lock granularity so when to put lock messages in the code which events should be locked and which not and we have the message semantic is the log message actually explaining the problem correctly or am I not able to get any information out of the log messages this is a typical thing that we see often the next part as mentioned as well all right studies how to gather these logs from the system so we have the scraping at the beginning so how to get locks from a software into our into our log management system we need to decide on scraping intervals how to filter them how to do the parsing right at the beginning probably how to do the batching in order to reduce the amount of requests to our to our lock um management system and also the contextualization this is a very important part because we will probably it's easy to gather lock messages from each individual service but how to make um some how to establish connections between the log messages in a distributed system this is a very uh very hard thing to do and it's challenging um and the last the the second part of course is the Lock Storage so how do we do we compress there do we do some kind of filtering and parsing there we probably need to do some format conversion because we always need to deal with some kind of Legacy systems in our uh in our environment there we need to convert older log formats to new ones and also to decide on the retention and finally after we establish these three steps we will end up with some kind of log analytics where we are able to analyze the log message to some kind of anomaly detection or machine learning model training to visualize causality and correlation between block messages and to be able finally to um to to gain the benefits out of it so basically reduce the mean time to repair so yeah I think we can agree that um log analytics the journey it's very extensive and comprehensive and and kind of difficult uh complex Journey so alexina sorry you know on the um the bit about the ingesting of data I was with a vendor um earlier earlier last week actually um because I was highlighting you the the costs of in ingesting terabytes of data into a log analytics platform obviously it's quite expensive maybe you know you get hit twice by you get hit by the the vendor that you want to ingest your logs to because they'll have a storage cost and you might also get hit by the egress charges from let's say Azure log analytics that you know they don't let you take the data out for free what these what some vendors are now looking at is how can they just um how can I create an API into the an existing log repository and just um view the logs and apply their calculations and metrics without actually having to ingest the logs into the data platform which I think is quite a good way it'd be a good selling point for um a log analytics company um so you know you don't need to stream the logs to or ingest the logs to your log analytics platform the log analytics platform can actually view them in the repository that they sit in for example it could be Azure log analytics could be AWS it could be mulesoft it could be wherever you you your your product is already shipping its logs to and I think that's a great way forward for it yeah that's a very that's a very good question so basically um all the customers that we are currently talking to are complaining or one point of complaining is there um is the cost basically right so how to decide which locks to forward which locks the store and um how long in the retention period should be um I would say there are two aspects um one aspect is um are there regulations to store log messages because if the company needs to comply with certain regulations then there's no way around storing most of the log messages for a certain amount of time right um we're working with a lot of customers that are not or they do not need to do it but still they're doing it because they're having this kind of fear of missing out so when something happens what happens if I don't have the required information this um and this can be tackled um or we are actually looking into that at the moment is how to filter log messages right at the beginning so storing only the log messages that are indeed relevant and reducing redundancy so there and the first step that we are that we're doing is basically to look at redundant information and um store only or store redundant information in a way that it's um it's not storing all the log messages but only the ones that contain uh the information once this is one part and the other part you mentioned as well is um what about the analytics um and there I would say it depends a little bit on the solution that uh that is applied when there are certain vendors out there they provide both right they provide Lock storage and lock analytics at the same time and there it's um I would say it's not too much of an issue you will store the log messages with these vendors you will do analytics inside of these systems and then you won't have additional costs of sending log message to certain to an external analytics system however there are also um the possibility where you for example store your locks in an S3 bucket and AWS right and then you would like to forward these log messages to a log analytics platform in order to gain insights and this is the case that you mentioned As I understood um there in my opinion or the approach that we thought about there is again some kind of filtering so without or to not have the requirement to send all the log messages from um an S3 buckets to the local analytics platform but only the ones that are relevant for a certain incident yep yeah but I think yeah yeah let's let's continue first yeah no actually um you already kind of kicked out of the panel discussion because I said the first question I or the topic scene we I wanted to ask was um how can we determine golden valuable lot data to ingest into the log analytics or the health lock analytics and data selection approaches but you all we kind of already touched upon that and because you also mentioned that you selected for example eight data sources and management storage and of course the more you choose then I want to ask like since you can how can we select the valuable log data and the sources and do you have your own maybe approaches developed kind of methodologies or Frameworks or do you have kind of rules or uh um I don't know to make the trade-off which data to select and ingest and which not I I think for us um it's it's we're probably not as advanced as selecting golden data valuable data um if I give one example um you know out of them 12 or 14 million log lines of data not all of them are going to be golden or valuable that get ingested every day um what because there's so much data there to identify what's golden and what's the most valuable is you know it's a big task and that's just on one system if we were to roll this out Enterprise why because we've only done it to one system at the minute it's an interesting thought how do you identify it and we haven't got a formula for it at the minute what we do is we work with the product team to say okay just just an example what API do you need to be made aware of and then we'll we'll put filters into our analytics engine to say if this API raises an event then you need to make sure that they get alerted everything else is just complementary I mean we even ingest informational uh logs as well just so we've got the complete picture um we don't have a formula Alex probably as um because he's far more knowledgeable on this than than what any of us are but yeah at the minute was I still call us we're an adolescent when it comes to log analytics you know we've gone past the the baby years we're not quite a mature adult just yet but we're going along the journey and any any help you've got on that Alex we appreciate it I mean Esther there are very interesting insights um whatever ask um is for example when you do this kind of analytics part um do you continuously analyze um your log message stream or do you do it on demand when um when let's say critical events occur or like you mentioned right when a certain API is raising an event then you start the log analytics uh yeah now it's continually running so in our platform it's obviously looking for anomalies from expected patterns so yeah yeah it's continually running 24 7 365. okay I see it needs some uh also days of data like of volumes of data to at some point produce good results because when it's the training period exactly I see I see okay because um yeah for for us um a draw maybe a slightly different picture here I'll give a few if you don't mind um what we're you know what we're usually implementing is um an approach that is uh not continuously analyzing log messages but we first of all Define some kind of golden signals for our customers so these are basically um these are basically um signals or metrics that are facing the customer and which are defined as slas or slis and when this is where we run an analysis 24 7 all the time so a certain endpoint for example in in needs to answer within uh 50 milliseconds so this this would be a very clear requirement and if it doesn't answer um then an alarm will will be raised whenever such um whenever violations of slas or SLI secure this is the point where we trigger the analytics of locks on demand so then we go into the Lock storage system and try to identify all the relevant log messages for the violation of a certain SLA and SLI and this is we do more sophisticated analysis than only on these logs and then present the results to the operations team or a certain uh developer team that are responsible for a certain service and then they can figure out what actually happened uh why this violation of an SLI on SLA appeared this is our approach maybe to um to have or my answer to this question is how to determine golden valuable log messages is it's about contextualization because log messages itself um they not always contain the information that is relevant for the product itself right the tests some kind of business value behind it so what will happen for example if there are log messages that should not be there I mean at the end if the customer is not affected I don't care too much um so it's fine they can be there um but again here um I would say it's it's the perspective of of us because uh we usually don't deal with very critical infrastructure what might be something different for you Ash when you deal with critical infrastructures that might be a different story with that approach Alex do you miss out on them early signals that might be um a warning that something is creepy if you're only going after the golden signals let's say let's just say it doesn't breach an SLI so let's say um a SQL Server normally has I don't know 500 connections to it and all of a sudden you only see on a particular time on a particular day there's only 50. now that might not breach an SLI but for us that's an early signal that something's not quite right you know there's no impact yet but we want that early signal that says beware guys this isn't what we normally see on this time of day go and have a look and it raises the alert um to whatever resolver team it needs before it even impacts on the application on the product um that's a very good point I mean if it's possible to identify such um patterns or such such signals very precisely then it's we can also or we tend to also set them up as early indicators um but here I would say there's a trade-off right between um specificity and generality right so do you want to take a look at everything and be notified about any deviations from a certain normal State yeah this tends to be um rather or more expensive and also sometimes error prone because you have a lot of false false alarms there as you mentioned as well during at the beginning um or do you know or you learned about your system in a way that you know the critical business processes and also critical execution paths through your distributed system and then you can set up these kind of probes that are monitoring the golden signals or early indicators as well and you will be only notified if they are they are violated right and but for the audience there's a q a box where you can submit your questions for the panel uh or your thoughts and ideas so we can have a good discussion and uh engaging discussion with you um but for the for this um question I don't see any questions from the audience and but the greater insights in data selection approaches like um really a great two no and hear about especially the early signals you mentioned like um creating patterns um that's the really the essence of the log analytics like it's really uh great to hear and no so I will then go to the our second question or the C for today because um uh as you oh I got actually um question from the audience right when I uh went to the next one um so the first question is if it's an application then log only from application would surface suffice or you have to take logs from application as well and the second questions may be more related to this question but yeah for the question from the audience it's whether only the lock is only from application would be enough or we'll just have to take logs from applications uh database Etc as well um I don't know if you want me to go first on give give our experience of what we do um so we take logs from the application the infrastructure the network any anything so we typically service map our applications and products and we understand the Upstream impacts and the downstream impact from a technology level technical component level so if there is a switch or a router or a database it might be a shared database that impacts our application of product we want the logs from that as well now the application team might not be interested in them logs uh to the level they would be their own ones but what we what we do is if the if that if that technical component is showing errors or Warnings or critical uh issues we want the application team to be aware of it as a as a customer not not as somebody's got to go and resolve anything and that's that's because we don't want it you know they might see us what I call a symptom in their application they haven't got an error well they've got an error but they haven't got um something to go and fix because the the issue is further Downstream on a shared resource we want them to know yes there is an there is an impact on your application you don't need to resolve anything so don't get a war room don't all jump onto the application trying to work out what the problem is we already know about it because we've took the logs from the other technical components and that resolver team is already dealing with it uh Alex I don't know if you want to add anything from your point of view yeah I mean you're you're totally right that I mean the challenge I would maybe highlight a specific challenge there um is to basically establish then at some point the correlation right you have the application locks on the application layer and you have uh the logs from properly managed systems like a database or in a message queue and then you have uh also probably locks from the servers as well from the infrastructure and now the question is how to make uh some kind of connection between let's say a certain execution inside of the application what are the logs that were produced when this was done what are the logs that were produced when a certain message at the end was written to the message queue what kind of logs were produced at that time um at the at the bottom at infrastructure I mean you need this kind of information as you said you need them in order to establish correlation causality in order to pinpoint then the root cause actually of a problem great um as we have great discussions time flies and I have to be also mindful of time um there's another question regarding this but maybe we can answer it at the end and let and we still have two kind of questions seems uh to go on let me uh ask you um this question like how does health log analytics impact the operating models of it organizations and this is more on once you implement like what Ash and uh mentions so our team has more time to innovate product and I will be very interested in um what the um what that would mean or how that goes and and also Alex you also said like increase of productivity so once uh log analytics is implemented so I want to delve more into that um and for that we have let's say five minutes kind of and there's also a question from the audience um yeah let me first ask you uh Ash and Alex how does it impact the it organizations and maybe even your organization I'll let you go first this time Alex all right yeah um I mean yes to make it short um or I try to make it short um usually what we can observe is that the team starts to start to work um more let's say agile or if there is a certain devops culture already implemented in the company then establishing log analytics um strengthened strengthens there this uh this this uh cultural aspect um because um at some point um when you take a look at the locks that are ingested in the lock analytic system they need to be written in a good way they need to be contain all the information that is actually needed later to identify the root cause of a problem and therefore there's a certain awareness between a developer that actually writes a log statement and then this developer will later see oh my locks X my locks are actually helping the team to pinpoint problems in the system which is um yeah which which closes I would say the Gap and lets the teams cooperate um more close closely together and the other part is Things become more agile right so when you have a structured log analytics you can establish also verification of your feature releases you can establish a verification of of your tests that are running in your in your infrastructure which will boost the confidence in releasing new things into your into your productive environment oh great so also more collaboration between teams you say yeah so I have a slightly different Twist on it um so um out impacts or operate if I think about um to start with there is a a significant amount of effort needed by the operations teams to understand all the logs significant mark them as significant so that when they do occur um you get an alert or an instant raise so there is a bit of training of the machine learning algorithm inside your analytics platform and that can take a little bit of time so people need to be able to people need to be aware that they need to put the effort in to train the system because you know an analytics platform isn't going to know what an API 50 millisecond means to whether it's normal or not they need to be trained on this is significant this is not significant etc etc if I think if I think a little bit further along that line what I would like to see happening you know it's a vision uh some of my colleagues inside National Grid also share is once we've got the system trained to know what's significant and what isn't then can we apply automation to going to enhance the alert by getting additional information doing self-healing remove that level one or tier one triage in that's normally done by a human and allow them people to potentially be trained to be a tier 2 or a tier three and go and build their career and gain their additional knowledge to better themselves the laborious tier one stuff we can we can use the automation tools to do that kind of um work and allow our colleagues and our teams to grow and enhance themselves and work on things that are probably more attractive rather than just going to get additional information for a log or an incident then you've become a much more leaner agile operations team because you've not got to have 50 people looking at alerts that are coming in you now have a a leaner more agile team of maybe 20 highly skilled productive um that can go and innovate the product they can spend more time engineering rather than just analyzing logs let the automation let the computer do the the boring work and then you you can build a really good strong team of um highly skilled people indeed um from the audience we have somebody asked so do we need to set up a dedicated Ops analytic team for fast that's a research so what would you say so I would say yet to start within in the early early part of it um you need the teams you need to have the team stood up to be able to put the effort into so the machine can understand what's significant and then isn't but after that so you know you you kind of want to get away from somebody sitting there looking at screens looking at dashboards waiting for things to go red what you want is Alex mentioned it earlier a devops team who are developing and enhancing the product and then respond to the operation side of it when an alert is made that's significant so I wouldn't say an Ops analytics team you just want a devops team who can do both Alex I don't know if you want to add anything to that no it summarize it perfectly I mean yeah as I said you need to you need some kind of transition period right to get your team from a current status to to um basically embrace the benefits of log analytics great great I hope that answers the question um I have to be mindful of time um as we are approaching the ants um so you know we had great discussions and let's say I'm really excited to implement log analytics solution or health log analytics from uh and certain start this journey in my organization so what would then be the optimal healthwork analytics journey and how can we built a great business case and roadmap and to be mindful of the time I would say can you briefly if maybe a strategy to pursue tips it guides way to the point um so if someone decides to um uh yeah during this journey follow this journey as you pursued and followed yeah if you don't mind if I start Joseph I try to again keep it short um what I observed when we when we talked with customers is um it is really helpful it's some point to be to be um precisely aware of what you would like to achieve when you establish log analytics right because when you take a look at different companies there are different objectives and so the first question would be is lock analytics even the right way to achieve these objectives um that's the first part the second part is if you are understanding the at the end the the values and the benefits that you would like to achieve with lock analytics um then you would need to lay out uh lay out the path towards it so which kind of log management as I mentioned right it's it's a boring thing to do usually for companies but it needs to be done it needs to be done in a good way in order to get the most out of the log analytics part whenever thinking about log analytics um you need to think about how to actually collect the data to process the data to make this data available without it it will be difficult to have a good analytics part and if you do all of these things right at the end it's then you can yeah draw all the benefits from Lock Analytics nice there won't be an end you'll always be on this journey yeah yeah and always be hopefully there won't be an end yeah for us you'll always be iterating you'll always be enhancing you'll always be um tweaking things there'll be new logs that need to be on board there'll be new dashboards that somebody asks for so you know it isn't a start and finish like a typical project is or a typical um program of work it's almost it's an operations function you've got to treat it as this is your day-to-day activity you're not always it's always going to be there there's always going to be something to do on the platform one one I know we're running out of time okay but I would say bear in mind your your security teams and your sock because you might already be collecting a lot of the data that you need for operational uh Analytics and um bear in mind that you at some point in your journey of Vlog analytics do you need to compile combine your operational data with your security data so don't always think you have to have two separate products if you can use the same product use the same product you can do different things with the data you can slice it put it to different domains you can make sure that people can only see what they need to see I would work with security teams to see how you can leverage what they're already getting yeah indeed so like it's very a good point you mentioned the security it's always good to have good communication uh very good friends with the security team as we had discovered through like uh the reason of uh for example with the servicenow uh Alton projects when implementing its uh security Factor was very dominant in that so yeah I could agree with it great um there is um is there any question from the audience um there is one uh very specific question maybe we can answer that when uh once we um we finish this discussions um but furthermore I don't see any other questions coming up from the audience I think we can wrap up today's discussions um yeah we were to go from here um well this panel accessibility panel is recorded so it will be published later on our YouTube channel and and yeah we you can find it on our website to also book a strategy session but I want to really thank you Alex Ash for um for you great discussions for the great insights you shared the knowledge you shared was really great to have you and host you uh and amazing to talk about um uh yeah the log analytics and I mentioned always like it's a journey because there's no ends this may be a beginning but um there might not be an end and but it was really great to have you uh thank you so much and uh thank you for the audience for participating in this panel and discussions and I hope you enjoyed it and I hope it was useful and and take the insights uh yeah so yeah thank you no thank you very much for having us it's been a pleasure yeah thank you Ash thank you Aki for having me it was great yeah when I was uh was not enough uh um who knows maybe uh at some point in a second a follow-up who knows let's see all right thank you all right thanks a lot thank you have a nice day bye-bye
https://www.youtube.com/watch?v=_LMiJnZ4dLE