Get started with Document Intelligence (AI Academy)
hello everybody and welcome to AI Academy as always before we get started I just want to share the Safe Harbor notice everything I'm going to share today is part of the product that's available today but in case we get to talk about items that belong to the road map and that are not available yet please refer to the safe harbor notice and check back with your accoun team before making purchasing decisions if you are new to a Academy what do we do in those acmis academies are content that we bring for you to bring you fresh ideas and better understanding and practical guidance of our products we recall the session and then post them on our YouTube channel for letter consumption but if you are here with us today it's a great opportunity to ask your questions and get those answers and for that we kindly ask you to use the Q&A panel on Zoom additional resources are available for you so you know that you might know that we have a community forum for AI and intelligence that we just renamed generative Ai and intelligence for this a quick link Snorks you can post your questions find previously answered questions as well as articles and content that we produce as a matter of fact here and few examples of content that are relevant for today's session so we have a document intelligence quick start guide on the community and a document intelligence FAQ with all the common asked question about document intelligence as I mentioned the session is recorded and then posted on our now Community YouTube channel our goals for today we'll start with a five minute overview and then go into more details for and inside of an instance for an exercise and will save some time for question and today we're going to talk about how to get started with document intelligence data extraction my name is l Sanchez I'm the outbound product manager for document intelligence at service now many business processes today still require retrieving information from documents which can be timec consuming and slow down the overall process in this session we'll learn how to quickly extract data from documents and streamline workflows automating your document extraction process can significantly enhance operational efficiency by Expediting forment minimize manual effort while reducing it expenses so let's look at a few example documents can be involved in a lot of different workflows and here are a few examples where document intelligence can help in the technology workflow space updating software entitlements from old forms on the employee workflow side think about HR onboarding you can quickly extract data from identity documents direct deposit of forms such as W4 and I9 so that the onboarding happens faster for customer and Industry workflows you can gather data from attachments so that your case are red more efficiently in the finance space quickly onboarding customer it's a process that's known as kyc know your customer with identity documents and tax forms extraction automate insurance form processing in the insurance space for healthare and public sector in The Med medical field there's a lot of forms that are still used and you can automate the processing of those of those forms in the public sector where we deal with a lot of certificate documents automating those certificate doent processing can save time for Erp invoice and purchase order is something that's extremely timeconsuming today and extracting information quickly from those can help speed up efficiency and finally any workflow that is based and built on the now platform and that involved document can be integrated with document intelligence for faster processing so all of these workflows can be achieved with document intelligence so that we accelerate the extraction of information from those structured and semi structured documents in turn we get reduce processing time we can minimize data entry errors and because it's based on AI it will learn intelligently so that it can adapt over time it's not because it involves AI that it's difficult to handle actually it's quite quick and easy to set up because it's all based on a low code configuration so no need for AI expertise because it's all part of the guided experience that AI learns over time so that it reduce the amount of work an agent has to do and increasing the document throughput and because of that by Design approach we provide Time Savings from day one and all of that as I mentioned is built on the now platform so you can expect all the benefit from that as well such as a secure AI pipeline seamless integration to our workflow engine and data sources and buil-in analytics if we say that differently we say we can move a process that is based on assigning documents and a lot of swiv sharing from Agents that has to move between different applications type data manually that they can read from a document and we can move that to a process where an end user can create a request easily via all the different channels available on the platform whether it's the portal the mobile virtual agent or an email and they can upload the document the document is then assigned to a case and rout it to an agent who can extract the value with the help of document intelligence once that's done the workflow is executed and the request is fulfilled that's how we get a streamlined workflow and a more efficient process so what how do we get started with document intelligence so here are the few steps so first we'll go in the instance and install the document intelligence admin plugin doing so also installs all the required dependencies then we want to gather and organize our sample documents we'll need at least 10 documents and it's 10 documents per each type and format so we want to organize them by type or by layout or by template we also want to take some time to review the fields to extract and maybe we need to bring in subject matter experts to make sure we understand the meaning of those fields then we're going to think about which integration method we are using so we can leverage our flow templates to build a custom app apption to manage that workflow or we can uh leverage one of the outof the Box integration for example customer service management the CSM uh integration Financial Services operation fso and account payables operation so you can learn more about our integration method in the docs and for the exercise today we'll use the flow templates after that we get to the fourth step which is to create create our document intelligence use case the document intelligence use case is the definition of our model and the fields to extract after that we can configure the flows as I mentioned or the outof the boox integration is that's if that's the way we are chose and then we can start processing a few documents the first few documents that we are processing are the initial training of the model and they are quite important and for the model to accurately understand what we are doing so this those first few documents will be done with a manual review after that we can assess the results and look at the accuracy of our current model and Define which level of automation we are going to enable for our flow and those could be recommendation which is a manual review autoi which is providing the values before review or finally full automation which is where the agent doesn't need to review values at all and all of that really depends on what your targeted outcomes are and the accuracy of the model based on the number of documents that trade so with that being said let's move to the exercise so as I mentioned the first step would be to install the plugin so we go to our brand new application manager type document intelligence and look for the document intelligence admin plugin and then we just install that one and it will install all the dependencies I I've done that prior to the session so that it's available and then we are looking at uh how we're going to integrate so for our case we actually have a table on our platform that we've built to capture the data that we want to extract if we would look at the documents we are going to extract today they are invoices and they look like that it's a service invoice with a some information on it and I need to extract some of those and so for that uh somebody in my organization built this application with that form here where I can attach my invoices and I want the extraction to be done so that the values are available I have the invoice date the invoice number the total of the invoice who's the supplier and then for each line on my table I have a service line item here and I will take note of the name of the table that I will reuse in my next step so once that's done and I have a good understanding of where my data is going to be going I can open document intelligence so I'll type document intelligence in my menu and I will navigate to document data extraction Administration and click on use cases this is the document intelligence admin this is where I can create my new use Case by clicking on new use case here and I'll give it a name so we saw those invoices were invoices for service so I'll name them service invoice and I will select my table for my data to be extract extracted to and I have to make sure to select the right table there so I will actually get the name of the table there we go all right so that's the service table here and I can save so that was my first step to create my use case I just created it um and right now it doesn't have any Fields yet so I as a next step I will will Define my fields and when I create a field I have different options I can create a single film which is essentially just a piece of data on the document itself I can create a list of checkboxes if I wanted to extract checkboxes I can extract a table which is a list of different items that all have the same columns but I can have multiple of them and the key here is that with a table I never know how many items I might have in my table and I can also create a single field Group when I want to gather Fields together so the good example for that is for example an address an address typically has a street a state city Etc so that could be grouped into a field so for us we'll have to start with an invoice date so I'll select single field for my invoice date I give it a name and then I can select the type the type is helpful for two different reason one uh when I select the values on the document as we'll see in the next step I can automatically make sure that the the the value is convert converted to the right meaning based on the type so for example for a date I'll select date and that way when I do my extraction I make sure that the the date is um properly properly converted and that's also used when I'm going to build my flow and I want to store that data into a field that is of type date on my table as well and that's what we do here when we select the Target Field so that's my invoice date field on the service table I want to keep building field so I'm going to check that box here so I can select and go go ahead and create my next field my next field is my invoice number and because my invoice number I I don't really know um the format it could have letters it could have numbers I'm just going to keep it as a text and I'll assign it to the invoice number field then I I have an invoice total which is the total value of my invoice and so for that I'm going to use decimal so because my number is going to be a number that might have decimal I'll select decimal and then I'll store that in the invoice total field and then finally as a single field I also have the supplier which is uh the company that sent me the invoice and that I'm going to keep as a text and I'll store that in the supplier I am done with the single Fields so I unchecked that box and I click on Save I can see all my Fields here I see them the name type Etc the next thing I want to do uh is that I wanted to extract the table on the document as well and so for that I'll I'll recreate a new table for extraction and my table here are my invoice items so I could see on my documents previously that I had different items that was that were part of the table and I want extract those and what I'm going to do is actually after I extract them I want to store them on a different table that is related to my main table because I'm going to create one record per item and so for that I'll select my my other table which is the service line table that I created here and I also created a reference field to my service table so that I know that I can identify each line and and tag them against my service invoice here and I also make that required I forgot to mention that required is used for two different purposes the first use of that is that on the UI and we'll see that for extraction required field are part of the required Tab and they help me make sure that I extract all the relevant information actually if I leave one of those blank I'll get a message saying that I'm missing required field that's one thing and the second uh way that's being used is for when we get to full Automation and we'll see that towards the end to for a document to be automated without validation the system is only going to look at required field so that if I have field that are not as CR IAL I don't have to wait for those to be extracted for to move to full automation okay so I created my table my invoice items now I'm going to create my different column and so for that table I'm only going to extract the name of the item as a text and I'm going to store that in the item column and then I'm also going to extract the line total the total value of my line because when I add them together they should add to the total of the invoice and I can save so just like that I am done with the definition of my fields and the the initial configuration of my use case the next step after that is to start processing a few documents to look at uh how it behaves so for that I'll create a new document task directly here from the doc document intelligence admin and I'm going to attach a document here let me copy that name so I can reuse it here and I can click add extraction after I'm I'm done doing that I'll have to wait a few minutes the way it's done um it's using our AI pipeline our secured AI pipeline so we request the prediction to be done to our AI model so that piece there takes a few minutes so that's what I have to wait and essentially I have to wait for the task to be processed and so again going back here to what I'm trying to do I'm trying to extract the information from my from my service invoice and directly have that stored on the the table itself and I did want to mention that my service table here is is is a t is a custom table so for this specific example I'm going the route of a custom application because I couldn't let's say I couldn't find uh an existing application on the platform that met my needs so I Built My Own application and I built my own table here but that could work with any any table all right so I waited a few minutes the task is processed now I can move on to manual review so if we remember uh in the different steps to get started I did mention that the first few documents have to be manually reviewed and we talked about the different level of automation because we just created the use case and it's brand new we are in the mode that we call recommendation so it means that the there is no values that are offered yet I'll just have to review a few documents first and so what we mean by recommendation is that the system doesn't know yet what I'm looking to extract so I have to guide to guide the AI towards what I want to achieve and here for example looking at invoice date I have to look at what it looks like on my screen and I'm going to start typing the value and then this the system is going to give me suggestions and and those suggestions also are um highlighted in the document so I can select the right one so here uh I select the date and remember when I created the field I also assigned the type of date that's what I can see here is that the the date is converted to a date format so that I make sure that the values is the value is correct and that obviously in that format is not as critical but when you have the format between months year days and year depending on which country you're located in you might have to pay extra attention here to make sure that the values are understood properly so that's for my invoice date then I'm going to do same thing for the number so I'm looking for that 1623 here on the document and I can select the value and we can also see here the confidence level on those values is 0% because it's the first time that my model is seeing that document so is not yet able to make prediction that will come after I've processed those few documents I'm extracting the invoice total and also here I want to show something is that on this document specific Al I have two values that are the same one is the total and one is the sub total and they might be the same number but they don't mean the same thing to me so I'm really looking for the total and not the sub total here and so because I have knowledge of the document that's why I can make sure that I select the right value and this is extremely important because if I select the sub total here I'm going to start training my model to do something that I don't want so I have to pay extra attention to the field I'm extracting because I'm teaching a model who's then going to be able to do that on its own and then I'll extract my supplier MC and then I'll move on to extract the table here so for the table I have it's that table here I'm looking to extract and for each line or for which row I'm going to extract the values so I have network security and then I have the line total then I can create a new row and go to the next line data back up and then my line total and then I can close that and then I'm done with all my extraction as I mentioned I'll go to the required tab here to make sure that all my required field are extracted before I can submit it I submit that and I can close this tab and now I refresh the list here I see the status of this task is done I'm going to to add another one so we can look at at the difference between the first one and the second one and then we'll move on to a model that I've processed more than those documents so that we'll see the the end result so I upload my second document and we'll see the difference of accuracy and so for now I'm just using the document task as a container for my document here I'm I'm doing that from the document intelligence admin but as we are as the next step we are going to build the flows using that Integrations tab here using the flow templates and once the integration the Integrations are going to be built I won't even have to upload the document on the task here directly because they are going to be attached to my record and and as the extraction is done the the flows are going to handle the part of creating a document task and then after the extraction is done copying the values over to my record so that's going to be handled by the flows that we're going to get where we're going to get to and I'm going to wait another a couple of minutes there we go so that's processed so that's my second task I can open it in document intelligence and so now I see that the the menu here the drop down opens because I have more suggestions and I see the the level here has increased so I see that if I want to extract the date from the document here the the AI found a date on my document that looks like a match with a 70% confidence so my extraction start starts to get a lot faster see that invoice number here I get more accurate predictions so for the for the total still same issue here I have to still be extremely careful about the value that I pick because for the AI model it's going to be a bit longer to understand that I have two values but they are not the same so I have to keep on training my model my supplier here it didn't see the value in my first few candidates so I'll start typing the value again and as a general good practice if I don't see the value I can just keep on typing it until I get to the value I want and usually it shows there and then finally I'll do the same with my table all right I'm going to stop here for that one so like I should have finished the table here but for the sake of time let's move on to how it looks like after I've trained the model on multiple um document so I've done that before the session so we don't have to spend too much time extracting values from document I did process a few documents here and what I'm going to do now is I'm going to set up my integration so what I want to do here is to create two different flows the first one is to create a task every time a new service invoice is created and so that's a type of process task I can I make sure that check create flow is checked here and I could add a condition if I wanted to only trigger that on some specific filters for my service invoice but here I'm going to leave it blank and then I can open my flow by going to flow designer I can review the flow but it's a template uh and I can just activate all I need to do here is activate it if I'm a bit more technical Savvy I can look at what are the different steps but here I'll just activate it so that's the first step that first flow so what does that mean it means that now every time I create a new service record and I attach a document to it I save that what it's going to do my flow is going to create a new document task if I refresh here I should have a new document task that was just created now so that works the second flow I want to create is the one that extract the values so what that means is that every time I either review the values of the document or it's actually fully automated extraction the values then are going to be copied over to my record so I create the integration here again it's based on the template I'll just have to open it in flow designer and activate it so now I can open open my task see actually that message means that the task is not processed yet so I get a message telling me that I have to wait another few minutes because my task is not ready for prediction so I'll wait and it's available and so now that when I look at my different candidates I see the level of accuracy here has greatly improved I'm I'm um more like towards the 90% for those values so what that means is that why don't I instead of having to pick every every value there why don't I start looking at saving more time and so to do that I'm going to my settings on my use case Here and Now what I'm going to do is I'm going to enable autofield what autofs mean is that if my accuracy is greater than 80% the value should be directly available on my um experience there so I only have to verify it with a glance of an eye instead of having to type the value so that's my first automation step is autoi and so after I've done that and I open my task here I'm seeing that all the values now are prefilled so that way I'm saving a lot more time on my extraction what I have to do is review them making sure that they are right by only clicking on them making sure that my table also is right and see I get all my table extracted here and I am done that's done that's submitted so now that's my my first way to save time and if we actually go back to this form here when we uploaded the document I see that the values are available on my phone so it means that as a as I build my my business process I can uh increase my I can keep on building on top of my my app here to do something with those values I can send that to a third party system with an API I can use that on the platform itself Etc and then the last step of automation if when I've processed enough document I'm ready to enable fully automated mod and so that means that no agent review is required and again I have to set my threshold here I think 80% show that it was working pretty well he was able to pick up the values but still being able to pick up accurately so between 80 and 90% is the recommendation um and so for me 80% would work here so I'll enable that mode I'll save and I'll go back to my list of services and I'm going to create a new one that's the last one for today bear with me there we go I attach my invoice and I can save that record I'm going back to my table here to my experience to look at the document task and I can see that a new one was created now if if I'm an admin I want to I might want to look at um go to document tasks here that might be a view that is more appropriate for an admin or somebody who understands the platform a little bit better but if I'm looking at the tasks that you know we just created now I could see that those test were created recently and the one I just created um again is being process so what I'm hoping now is that all the values in the document reach a the appropriate threshold so the extraction is going to be done automatically and I don't have to review it all right so my status now is true is processed and this the is processed s is true and the status is done and if I go back to that page here I can actually see that this task was is straight through process meaning that it was actually picked up by the system accurately and the extraction was automated I can reload that form so I see my service lines there but what I'm seeing here is that I uploaded the document and now my AI is trained enough to do the extraction on its own so I didn't have to even review the values so if you think about the process how it how it was before I implemented that think about the the person who had to look at the document and type all those values manually and input them into a third party system well now all they have to do is upload that document on the record and the extraction is done automatically so that's a huge save of time it also minimizes the risk of Errors so I'll save some sometimes for questions someone is asking what is a document task for and I think we've understood that at first my document task is kind of used as a container to host my document but really in the in the back end it's it's the the mechanism that triggers the prediction so a document task should be a record that create for each individual document doent from which I want to extract my values the question that would come often is how many documents do we have to train my model on and that really is going to depend on the number of fields that you have the quality of the document the number of different layouts and Target outcome based on the level of accuracy and the best way to to look at it is to upload a document look at the results look at the the accuracy confidence level and keep doing that until you see the levels rise to the level that you think is a good way to automate so we definitely suggest at least starting with 10 documents and after that you can decide whether you can move to Auto or if you are uh accurate enough to go to full automation another thing I want to mention is that um the way an AI model works is that it's learning from different variation of the document so we don't recommend doing that task with the same document so if I if you saw that really well in there but all the documents that I'm using here they all different they all have different values if I were to do this task with the same document over and over I'm actually going to mistra my model and that's not good outcomes for the future so really make sure you have a a good samples of real documents to work with uh now who's doing that um that part here that we saw today um I think we saw that the experience is quite easy to manage so I say you don't need to be an advanced system admin to do that you probably need an admin to build the application or or or the rest of it but I think a business person could create their use case and the thing that they could bring is that as a business person they also understand um which fields are important in in the document so I think it's a combination between it's a it's a good collaboration between an admin to provide technical guidance and a business person to provide knowledge about the document somebody asked uh about the flow so we it was fairly easy to create the flow then we have to go there and activate them and essentially the question is why don't we just create them automatically but the the reason for that is that I if you remember at the beginning I mentioned that document intelligence could be integrated with other applications like account payable operations and in those cases you won't need the flow because the flow already exist as part of an existing application so we did want to give the freedom to choose between the two there what is controlling the processing time I think we a few times during the exercise we had to wait for the document to be processed that's uh based on our backend infrastructure and so it has to do with the type of Hardware we have in the in the in that back end um where that's located and and all of that but you know service now has been investing a lot in um the new generation of a of AI so we'll see that um it it it's only going to go down so it's only going to be faster and faster to do that as we invest more and more in our Hardware what plugin do you need to use document intelligence you have to go to the plug-in page and look for document intelligence admin and you install that one it's going to install all the different dependencies so you have a couple of additional plugins that are part of that but document intelligence admin can install all of them they will install the document intelligence plug-in and a few others as somebody mentioned something about migrating that's a great point so here when I'm comfortable with my model and I like the way it behaves what I can do here is come on my menu and click here and I click on add to update set and essentially I'm going to capture that use case all the field and all the training that happen and once that's done I can take my update set and move it to a different environment and uh we're going to finish with a few questions regarding licensing and then that's going to be it for today document intelligence is a paid plugin so you'll have to talk to your account team to know how to get access to that uh and there is a few variation whether you already own something like CSM Pro or not essentially it's it's part of automation engine but it's a paid plugin so talk to your account team for that all right with that we're going to conclude the session today thank you all for joining and see you next time
https://www.youtube.com/watch?v=sr6rUoVC480