Virtual Agent Academy: Improve NLU model performance with new feedback and testing tools
all right folks let's get started with introductions once again welcome everyone to virtual agent academy so let me just kick off with our additional resources so um as always we take a look at our virtual agent nlu community forum there you'll find all the latest guides information and as well as the ability to ask questions and give answers you can visit the link there at bitly v-a-n-l-u-com community for short if you're watching on youtube thanks for watching uh we have a whole playlist on virtual agent academy where you can find all the recordings of our previous virtual agent academy so yes this has been recorded and if you're watching and you like what you're seeing go ahead and click that like or subscribe button down below you'll get notified every time our recording becomes available recording for this video should be available uh within the next couple days certainly within the next week with that let's get started with today's topic today is a special topic we're going to learn um on about a san diego feature for nlu and you'll learn how to improve your nlu model performance with new feedback and testing tools and here to present that is our returning champion nelom when it comes to all our uh or many of our uh virtual agent academy and you topics it's it's it's on the forefront uh driving that so kudos and much gratitude to her i'm going to turn over to her now to go over our agenda and our new topic milima take it away thank you thank you victor let me go ahead and share my screen and welcome everyone to this session of uh nlu tuning um and uh the new features that are available in san diego so so we'll go over a quick overview uh for our goals for today uh and the topics that we will cover followed by an exercise where we will do a hands-on demo on the features we are covering and followed by q a session uh please feel free to post your questions in the chat and there will be folks that will be answering them as we go so starting with san diego we have introduced the expert feedback loop uh expert feedback loop is basically it uses your chat log to which is when we say chat log we mean the open nlu predict intent feedback table and it uses the chat log to extract a subset of the chat in data that is available in the logs uh that represents your model you know adequately for each intent and it provides an opportunity for your experts you know your domain experts your nlu admin to then go in and label that data uh to provide uh feedback to the model um back uh in terms of you know whether the predictions are correct incorrect uh or uh you know it shouldn't be predicting anything and we will go over the demo uh but the good thing about expert feedback loop is you know it kind of you don't have to go through a whole big set of data uh you can look at the extracted uh data that the system provides using expert feedback loop and then label those that data um this chat log data obviously is you know typically available in the production right where users are requesting things uh there are a few ways to go about it you know typically we uh would not you know tune the model directly uh in production you know you may want to try it out in a subprod environment so that you can basically you know export xml uh for your for that table into your subroad and then do the tuning there and then once you are comfortable you can move the changes to production right and the system expects a default of a minimum of at least 300 uh so this uh basically the expert feedback loop runs periodically incrementally so you know the system default is 300 uh chat blocks and samples for the expert feedback loop to contain um data to label uh and then once you have uh labeled hundred samples at least then the tuning opportunity is available and these are system properties you know so you can adjust them based on your need uh this is a screenshot of um expert feedback loop and we will go over the uh demo portion of it you know uh but basically you know whichever models are available for expert feedback loop those will be available in the top drop down and then the intents are listed over here and on the left bar and then you can click on each intent to do the tuning and labeling and in san diego we also have the test panel feedback feature where you know you can use the try model button uh within nlu workbench uh manage model phase uh to provide feedback to the system back uh based on the predictions that nlu is providing uh whether those predictions are correct uh or they are incorrect if you don't agree with them then the system will allow you to either pre pick a different uh prediction i'm sorry pick a different intent uh or if it if you think it goes to no prediction then you know uh you can select that as well uh and then once you train the model then that information is used to uh tune the model further you know and uh the system would expect you know um a few um [Music] a few some like you know um quite a few a few samples to be um [Music] provided the feedback to be provided uh and then train for there to be a change in the model performance uh some of the other uh enhancements for nlu in san diego release that you uh can be aware of i mean expert feedback loop and does panel feedback we are covering in this session uh with model creation we've already also introduced uh uh importing from csv um and in fact the model that i will be uh you know demoing in the demo this is what we used uh and then you can also export nlu models to csv which is also very handy feature in san diego we also introduced adding nlu models to update sets right from the cis nlu model record from the model table record as a it's a related ui action on the form and in san diego we have now have the phased model building approach which means that you know you have different it has separated the process so making it a little bit more intuitive uh you know where when you're creating the model you go into manage model and we'll go over it during the demo as well are you going to manage your model and then the testing phase and followed by the publishing phase uh san diego also workbench also has the resolve intent issues feature which is pretty handy while you're building your model if the system finds that there is a conflict in your intents uh it gives you that feedback right away in workbench and there you can you know resolve those conflicts within workbench itself so now coming to the exercise part um sorry okay can everyone see my screen yep okay great thank you thanks victor uh okay so here i just i just came to the instance and i'm in workbench right now and as we can see some of the features i talked about are here uh using import data from csv uh you know and this is the model that we will be working on today um it's a english only model and we'll just look at the intents quickly so we are familiar with uh what the model contains it's a combination of uh hr and i t related uh requests um you know as we see there is powerpoint template and octa reset uh part of ipsn and then uh we have some hr related ones as well like compliance policy and trading window info etc and then there is also some setup topic so it's a combined model okay so now on the other tab i have uh this model is already uh created trained and published on my instance and i ran one batch test against the model uh and so we uh it's basically you know currently it's predicting at 89.79 percent uh with about 23 or so uh incorrect predictions uh and the threshold is 60 sec to 60 so that's what we are seeing on the right side but for the purpose of this demo we will focus on the uh model performance on the left over here right ah so and how we can use expert feedback loop to improve the model performance so now coming back to um expert feedback loop um i will go ahead and open the module so in this case in this instance i have enough data that uh the system um you know was able to extract from the chat logs so we are seeing there are about 70 to 79 um samples per intent in the model uh that are available for labeling and uh you know like i said uh before like you know you don't need to do all of them uh but there is a at least you know you need to uh make sure you're covering uh enough samples uh within each intent uh out of the box you know we need a hundred labels obviously for this demo we will not do hundred uh but if you know the we have 11 intents and then you know you want to make sure you have labeled uh or each of these intents um you know proportionately when you're doing the labeling uh for the purpose of this demo we will just go ahead and do a few and then see how it changes the model performance right so here in this instance i have only one model that has expert feedback loop data so that's all we see otherwise if you had multiple models that have uh labeling opportunity for expert feedback loop you would see that in the drop down um and then you know you can click on each of these indents to see what the samples for labeling that are available i'll go ahead and start with admin access and for the purpose of this demo i'm focusing on some samples uh that are not just uh incorrectly predicting as we can see over here uh but there were some samples that i saw that uh should have not made any predictions you know so i'll go ahead and label those for this demo and then we can tune the model after the labeling and see what the difference is so here i've seen there is a mfa reset issue that in this case this is octa related so i'll go ahead and label it as doctor [Music] and i'll go ahead and save my feedback so coming into compliance policy uh for compliance policy i saw there were some samples around powerpoint um yeah so this is a sample ppt deck but i do have a powerpoint template intent uh but as we can see you know there's not enough information in this uh sample that you know should it should go to powerpoint template uh i would expect this to go to search fallback because of no context you know so i'll go ahead and say mismatch this does not belong to compliance policy and i will mark it as not relevant to any intent and if i try to go out of this intent the system will uh alert me and you know i do need to save all the feedback before i go to the next intent now in the case that you're not sure about uh some uh samples right like for example maybe bank financial corporation what uh the prediction should be you can mark it as unsure and what it will what the system will do is you know it will uh place it in needs further review um after this is saved so now i have one done for this sample and another one is needs for the review which is the what i just labeled um coming into uh one more sample let's do from this uh intent uh i did see there was a pr yeah so this one again is not relevant to any of the intents so i'll mark it as not relevant and i will see if um next let's try octa reset uh octory settings here for octa reset there were some predictions for the word university i guess not uh let me try i thought there was something about sales request yeah so sales request this has nothing to do with octa reset so i will go ahead and again there is not enough context in this sample either so i will go ahead and mark it as not relevant and i saw something compliance related as well yes so this one again is a mismatch but there's not really anything about policy you know it's just the word compliance you know so i would not want to direct this one to um compliance policy you know instead i'm going to mark it as irrelevant and i will save the form and save the page um let's do one more and then i can go ahead and tune the model i'll try um trading window on trading when i thought i had okay i'm right yeah so here there was something on sso issues this one is a mismatch and i can mark it as octa reset as the correct intent and i thought there was one more um powerpoint yeah there was one on powerpoint okay so this one i'll go back to to do yeah so this one obviously goes to powerpoint template so i will mark that as the intent and i'll go ahead and save my uh page and do one more purchase requisition so this one again it's not a relevant intent and i'm going to mark it as such and then one more intuition um there was one more powerpoint unless i already did that but no okay okay so now i think uh yeah and so i do need to do 10 anyways so i'm not at 10 yet this instance is set to do 10 let's see who are you this one i'll mark as also not relevant now save the page uh this is obviously a not correct um intent access p slip uh i don't think i have any available intent i will not this is not relevant um can i trade my stock should be trading window okay now action this is very strange i don't know why it is expecting i know in this instance it was um set to 100 i'm sorry set to 10 so let's see if i can remember this property okay maybe i missed something i missed hold on yeah so this is a good exercise for us to cover do you want the uh workbench glide optim optimize min was that the one you're looking for uh optimize yes this one let's see is it i'm gonna paste it in chat okay uh this one let me see [Music] i think so yes thank you victor no worries okay come back to our system property i had to switch instances in the last minute i was doing okay there this is hundred so i will switch it to 10 and i will say ignore cash and save the form okay so now once i reload the page i would expect enhance the set to be available oh so i'll go ahead and select my test set so at this point we had already labeled uh 10 samples so and we change the system properties so that's good we got to see that also in the demo and i'll go ahead and submit uh the labeling for optimize so while this is running um you know i will not be able to go to workbench to show the test panel feedback so what i'll do is i'll go ahead and use this other instance that i have and open mmu and we can review the test panel feedback feature and then come back and see the model performance from uh expert feedback loop so i have the same model on this instance as well with the same set of intents etc and we can use the try model feature to provide feedback um so i'll go ahead and maybe pick on the first intent over here new logo and let's see what it is predicting uh and predicted with 90 so that's great i will give feedback that this is the correct prediction um then let's do just a sample uh powerpoint uh and it matched powerpoint template with hundred percent confidence because of the word powerpoint now like we were uh seeing uh before right just the word powerpoint i don't want it to go to this intent you know it should be if it is a template related uh intent right so i'll go ahead and say this is not the right prediction no intense save changes so once i do enough of these then you know it will go uh it will tune the model accordingly uh let's do one more purchase requisition we can pick on okay so here this already predicted purchase requisition right so that is good it picked up on the word rec i didn't uh i didn't uh say requisition and i don't have any vocabulary set up but the system did pick it up so which is good so that's basically the uh tri-model feature and one more thing like if it predicted wrong i could have also let's do that now mobile app um okay so it looks like it is predicting now mobile app but let's say it did not predict now mobile app um i could have you know said this is the correct prediction is now mobile app in this case but it did predict uh now mobile app so obviously it's not showing up in the list over here uh so now so that's the uh tri-model uh feature with feedback and here once you have done enough uh samples uh testing and provided the feedback when you train the model uh and then retest it using batch testing you should see that a change in the performance uh coming to oh wow okay nice everything was working fine huh okay so i guess we there is an error that is coming up let me see if i can show the um feedback that we did on this instance so i hope that will be available okay so here on this instance uh i had done a dry run right before uh the demo and this is already published and available uh it's basically the sampling was almost the same i did a few of little fewer samples in this instance than what was done on the other instance so here so this is what the ui would look like after the expert feedback loop you know model tuning completes so as we saw when we started right the model was performing at 90.64 accuracy with about 21 incorrect and the labeling uh that we did the system picked up you know fewer incorrects based on the [Music] feedback that was provided and then you know basically once the publish optimized model um dialog is available uh you will be able to publish the model over here so with that we will conclude the uh demo uh part of the uh session um and i think we have a few uh questions also right and we are almost up for time anyways so [Music] yeah so let's go ahead and get to the questions uh nilma there's a great list of questions and i want to make sure we read these to get captured on audio um first question is do you have any recommendations should customers do this in production or sub prod and are there pros and cons to each correct thank you that's a great question and uh absolutely we would recommend you know uh tuning the model in support first to see uh you know what the impact is you could do it even in production uh because as we saw right you it doesn't automatically publish the model straight away and you can you know choose not to publish it so that is an option but the platform allows like different options for um doing the tuning in sub broad environment you know so you like i have mentioned like we can export xml of the data uh that you want to uh use in subpro so you can take that data from the table um from your production instance bring it over using import xml into your support instance and then do the tuning over there and excuse me once the tuning is done then you can move your model by update sets you know to production um you do have the option to do it directly in production as well i mean it's not you know you if if the results are not what they need to be uh you just won't publish the model right and the way the tuning optimization works is you know uh it takes into effect only until the model is published so once you make any change to your model uh you know in the train samples as well for example right uh you do need to you know then uh test and tune you know do the batch testing and publishing uh again right so do you to the question you know it just depends on what the preference is but even if you were to do it in production it's not going to you know hurt you know it's not a big impact as long as you don't publish it if it is not giving you the desired results but ideally you want uh perform these steps in some production first yeah and that's a great call out to leverage the new capability to upload via csv into the nlu workbench so that one's a new new capability for san diego um so a little bit more on this new nlu editor role um because again i think this is one of the the first opportunities to introduce that role and what is required let's say if you want to bring in a subject matter expert to help with the expert feedback loop um do you recommend they have the admin role or just the editor role to go ahead and label or categorize what should be in the model yeah it all depends on what the expertises on the person that would be labeling uh ideally you know if they're not familiar with workbench then you know they could be assigned just the nlu editor role and then you know they can do the labeling and then the actual nlu admin can take care of you know publishing the model and all of that um so yeah okay and last question is are these capabilities available for other languages and in the ones mentioned are brazilian portuguese spanish english um where do we stand with that oh that is something [Music] as far as i know they are available but we can circle back and confirm and then um you know confirm that they are available for other languages as far as i know they are it is available for all the supported languages but let's let's wait to confirm it great and we'll we'll post that to the community and just uh um a final note or a call out there so for folks using brazilian portuguese the recommendation today is to leverage the new portuguese language available in the san diego release so that is guidance on if you want to leverage this for portuguese and we will follow up on the community on language support for some of the expert feedback loop and the uh the testing feature leveraging continual learning yep and even though for the previous question about editor versus admin we will also we will just make sure we confirm we have the right answer but you know admin role will definitely give uh full permission uh and you know labeling should be available to edit at home as well great well no further questions so thank you again everyone yes and thank you everyone for attending our va academy we'll see you again in two weeks be sure to fill out the poll if you haven't already we just love to learn how you are all doing on your nlu and virtual agent journey seems like um the majority of you are saying yes or i plan to which is great once again nlu is a great way to personalize and make your verse your legion that much more intelligent um in terms of using our expert feedback loop we understand that uh you know not many of you using it but hey you just learned about it today in academy and that's the first step so feel free to take a look again the recording is going to go online um in the few days after and once again thank you for joining us thank you thank you everyone
https://www.youtube.com/watch?v=aEmrup6u90A