What’s new in Vancouver for Document Intelligence (AI Academy)
so again uh welcome to AI Academy before we start I'd like to share our Safe Harbor notice everything we're going to present today we're going to talk about what's new in Vancouver for document intelligence everything we're going to share today is available today in the product it was released as part of the August and September stories but in case you have any questions regarding the roadmap and I do mention things that are coming in the roadmap in that case you'll have to apply our Safe Harbor notice and check with your account team before making purchasing decisions if you are new to Academy here is a little bit of what we do so this is first and foremost for you we bring you fresh ideas what's new in our products as well as providing you with a better understanding and practical guidance for our product this session is recorded and will be posted publicly for you to be available afterwards but if you are here with us today it's a great opportunity to ask your questions so for that I'll ask you I'd like to ask you to use the Q a panel in in the zoom interface to ask your questions and we'll make sure to get that answer my name is loik Sanchez and I'm an outbound product manager on our platform AI team and I'm actually joined today by Eliza which I'm I want to thank very much for joining me and helping me with the session today um in in addition to hosting our AI Academy sessions I also make sure that our Ai and intelligence Community from is running smoothly so for that there's a single links very simple link to access SN dot Works slash Ai and we did just post a lot of new resources regarding generative AI so I highly encourage you to go check out the AI Community Forum this is also the best place for you to ask your questions we have experts monitoring The Forum they can help you with the questions and you can also look at the previous previously asked and answer questions on and the community as I mentioned this session is recorded and then posted on our now Community YouTube channel now I also started sharing a bit of our product updates on LinkedIn so if you wanted to get a fresh and more direct access to our updates also feel free to follow me on LinkedIn under Lois Sanchez all right today we are going to talk about document intelligence and specifically what's new in Vancouver so as always in those sessions we try to keep the amount of slides minimal and spend most of the time on an instance so today I'll go briefly through an overview of what's new and then we'll spend most of the session in the instance actually doing some configuration and are always enough time for questions and before we start let me actually launch a quick poll so I would like to know we're going to talk about new features in document intelligence and specifically document classifier but before we talk about classifier I would like to get an understanding about whether you've heard about document data extraction which was our main feature of document intelligence when the community regions was released a few years ago so let me know if you are already familiar with um document intelligence data extraction and also how familiar are you all right I'm going to close the poll it looks like almost 50 50. so 42 percent of you have heard of it and 55 50 54 of you have not heard of it and it's smaller amount so only four percent uh actually testing it it's good very good to know thank you all right so today we're going to talk about document intelligence and what came out for Vancouver so document agents is actually a store application so today we're going to talk about specifically the version that came out in August and it's September so that's version 3.0 and 3.1 which is our latest version for document intelligence and on top of that we have an admin experience that helps you create your models and that's document intelligence admin and so that let us release for admin is version 2.1 now I do want to mention that both of those applications require an instance in Vancouver um if you are not in Vancouver you can use version two so if you're still in Utah you can still use document intelligence version 2.4 but today let's talk about what's new in Vancouver so the first thing is our new main capability which is called document classifier so what it is it's building upon the use of AI in in that case not for data extraction like our main capabilities but for classifying incoming documents so when we what we mean by classifying is actually finding the category of finding the type of a document and that's very helpful so that when you get different documents you can categorize them and then route them to the appropriate processing workflow now it could be routed for approval to a specific approval queue or it could be anything that can be done on the platform with product designer and all of these platform tools tool as well as routing that to the right extraction model so it's document extraction you can proceed with the extraction and that help the classifier helps you identify what type of document you are looking at so it's using AI so it's the same concept as document extraction in the sense that it starts with a phase of training and model and then using that guided experience to you know start saving time by looking at recommendation based on your previous data and then all the way up to full automation once you have enough data and as always it's we provide a experience for you to configure the different categories we we know everybody has a different way to look at categories they have different documents so really it's uh we are allowing you to oh we are helping you define your your categories based on what makes sense for you and it's an experience that familiar with the data extraction so it's there is no there's a unified extra experience between the two so that your agents are not lost here so let me just give you a quick example of that what that would look like so let's say for example you receive two documents from one of your uh input on their platform email record producer things like that and then those documents can can be sent to document classifier to be classified or could be categorized based on the machine learning or AI engine in the background we are able to identify those documents one of them is a receipt one of them is a certificate of proof of completion now that we have the information and that's pretty much automated we can route the receipt for extraction via a document extraction use case and the proof of conviction is just sent for approval and then the process can continue the payment can be processed and the reimbursement for the employee in that case for example can be done and then the workflow is closed so it's really showing how uh that those categies can integrate seamlessly into your existing or or new processes and workflows on the platform so how to get started with document classifier so there's a few easy steps to get started and I'll show you that in a second so first of all we create a document classification use case then we create the different fields for our categories so we're going to create one field per category and then we create our first document tasks so we attach different documents to our document tasks and we label the document so we proceed with labeling those documents and then once we have enough data we are ready to start training the use case now that our model is trained any incoming documents will be processed via the classifier and the prediction will be made on the category of that the that document so for the exercise today we're going to go in the instance and I'm going to use that scenario where I have a an input process in which I have three different documents one of them is a survey form from um Finance institution then I have a stock transfer request and then finally an authorization for a bank uh Bank bank transfer so let me go in my instance now and let's start with again saying that I am in my instance I have the latest version of document intelligence here 3.1 and admin 2.1 so that's what a way you're showing today so let's do classification so I'll start by typing document intelligence and I see in the menu now I have document classification and I can go to my use cases and create a new use case so I'll give that use case a name and as we saw we were dealing with Finance forms so I'm gonna name that Finance Finance forms and submit that and then I can open it now our second step was to create our Fields so our fields are our categories I create a first category for my survey so it was a survey with the checkbox and all I need is the name here and I click I'm going to create a new field so a second field for my it was a stock transfer request so it was a form to request the transfer of my stock so the stock transfer request and submit that and I add a third category and that was a Bank authorization form all right I created my three different categories and I'm ready to create my first task now so in the first step here of what we call the initial training I I'm creating if you tasks to tasks to do my my labeling so I create a first task with a set of documents and I batch them together there is a limitation on 10 attachments per task so I'll process them 10 by 10. and I want to make sure I have a data set with enough uh of each of the documents and we are aiming to have um a an overall 50 documents of each category so that reflects at a good balance data set so I attach my documents and then I can save my document task and then that's when I'm I'm going to click show and lock Intel here to see the the document in our document intelligence workspace and do again well like would call the labeling so there is no recommendation yet yes it's expected because we are doing the initial training so I click OK here and I can stop we also see there is again no prediction the model is not trained yet to know what we are looking at so we do that manual labeling so for each document our assign a category so my first document here is a survey so I'll click survey and I can move to my next document as soon as I click on that form on that field here is going to display that document on the screen here so I can easily do my labeling or categorization next one is also a survey and then after that I do have a different one so that's a stock transfer request and I can click here so here all my documents are single page but if they add multiple pages I would get an option to label each Pages differently so if if a document was a mix of different sub documents I could click on one of the options here to label each Pages differently but here they're all single pages so I'll continue with my labeling and I have my third type of document here foreign once I'm done here I can click on submit and I close that tab if I refresh at least I see that each document that I submitted was labeled here I see the the value that was assigned to them and I can navigate back to my use case so I would have to repeat that step a few times until I have enough documents and then once that's done I can click on train the use case so I am back under the use case form here and again once I've processed enough document I can click on that button and the training is going to stop so here as I mentioned I don't have enough documents so the information message here informs me that I might not have a great accuracy results in that training now I can take a few hours to process based on the number of documents so to save us some times I did a processed more documents in that um use case here and the train it so that we don't have to wait for it to complete so once a model is trained if I create a new document and assign it new documents that the model has not seen before so in that case I create a task like assign some documents but because my model was trained now I see a new button here a new action that is process task so I can click on that and in that case now the model is going to predict the category of the documents so again to save some time I just did that before and I'll open it now and I see that this first document here now is because of my previous data is being predicted to be a stock transfer request with a 47 confidence that's uh here helping me label that faster that second one the confidence level is not as good I will need more data and the last one 35 of being the US person form survey so that's uh here so you can do that with enough documents and then um once you are uh confidence in with the results you can as you could do previously with extractions.assigning autofill to the model saying that the the values would be directly filled in in the the workspace so we all start saving more and more time as we go all right so that was document classifier so there were a few other updates in Vancouver so for data extraction we are introducing the notion of required fields so when you create a data extraction model and you create your fields which are the values that you need to extract from your document you can specify whether those fields are required or not if they are required you'll get in message if you try to submit the task without Computing all the required fields and you also get another tab on the extraction workspace to to directly look at your required fields and when you enable full automation so that's what we used to call straight through processing it's the ability to extract values without any manual validation when you enable full automation mode and you have fields that are required and some that are not required only the required Fields would be considered for a full automation so that allows you to have a little more flexibility around what you do for for validation all right as part of the document intelligence admin experience um you we also improved the experience to create Fields so if you if you are I'm going to show you that in a second when you create a new field you can decide immediately between a single field a checkbox list a table or a field Group and when you select for example a table then you all have the different columns here that you can select from the same from the same pop-up and that's pretty convenient it makes it faster to create your use cases foreign we can also assign a type to data now so um previously pretty much everything was extracted as text and then post processing was happening to assign to type or to normalize the value the values for for too much which maybe a match either like the data or things like that so now it's possible to define a type so when you create a field you can assign a type of text or it could be a date or a number so intake a round number or decimal number float Etc and you can do that in the during the model operation phase and what that gives you is that during the validation of the extraction then you'll see a bit more guidance around converting that value so when I extract the date for example it's going to convert that into a universal format and I can modify that value if in if in fact the the numbers where if the date format was actually a bit different in my document I could go and change that and because of that then I can directly store the value as a date in my target table so I extract my date and then I store it as a date on my record on my table so that's again saving some time and providing more uh automation uh also as part of that document intelligence admin experience we are able to quickly duplicate a use case as well as export it or add it to an update set so that it could be moved between instances um so that's simplified now and finally uh the experience to extract tables as some new improvements so let me just show you we've seen enough slides let's go back to the instance all right so now I'm typing document intelligence again and in my menu so I just showed you previously the classification and now I have data extraction and data extraction admin so I'm going to go to admin and I'll show you uh what I've just mentioned as our announcement here so I land on the dock Intel admin experience okay all right and I have a I'm gonna take one example here so as part of that experience so as I mentioned we can duplicate that use case so that's now super simple in case you want to try that model on a different set of documents or if you wanted to expand your model to different types of templates you could do that here by duplicating the use case you get the use case all the fields and all the training that happened and you can also export that model so you can add it to an update set so it can be moved between instances very easily now when I create new fields uh so now that that experience here is new now I can add more guidance around what I'm trying to achieve so if I'm extracting a single piece of information from a document I can create a single field and I see here I can select the type of it so I can do it uh yeah it's a piece of text or is it date or a number and if I do want to create a table like I did for that one here I create a table I have directing my target table in my parent mapping field that I can select here I can make that table required or not and then all I have to do is create different columns here so it's pretty straightforward to create the different columns for my tables very very fast now and so uh when I do actually process a document now let me show you so we had one so that model is pretty well trained now so it gets saved very well very good recommendation it's autofill and I had one field here that's a date so that that field is set as a date it does the extraction from the document and then here it does the conversion so is the actual is there actually is it actually August 3rd or is my document in a different format in that case I can change that here if I want it all right now table so that experience to extract table has been improved here I also wanted to show you sorry the required so if I wanted to save time or only look at my required Fields I could just go to the tab here and bypass for example the cons the customer here was not a required field so you're going back to the table I can do my table extraction um I can so there's a few improvements here on the experience I can resize my columns and then I can clear the values add rows before or below that's a few new features so that makes it a bit easier and in that case I can also directly review that whole row here all right that's that's it for our overview and there's also some additional language support if you if you know until now it was available for English French German Spanish Portuguese Dutch and Italian and with version 3.1 we added a support for check Danish finish Norwegian and Swedish all right I have a quick final Poll for you and then I'll answer your questions so after seeing the content today would you start using document intelligence we showed classification which is our new feature we had previously a data extraction feature so I'm curious to know whether you that's helpful that would be helpful for for your business and for your workflows uh whether it's more towards classification or data extraction or both so let us know in the poll in the meantime there is a question is this part of a platform feature so is there an additional cost involved to leverage this feature doc Intel document intelligence in itself is a separate license that you need to to get via your account team if the question so once you have documentations you can use extraction or classification that's not there's no two separate licenses for that but so you do need a license for document intelligence uh that's something you would get if you are a a Pro customer and you can reach out to your account team and it's also uh so it's available mainly via automation engine but it's also embedded is some of our other pro products so it's part of CSM Pro for example or fso which is financial service operations and there is an APO if you heard about our account payable operations product that was released also recently it's also leveraging document intelligence so you can get document agents via those products as well but as always it's something you want to reach out to your account team to make sure you get a cliche and the answer that really works for your specific use case all right I'm going to close the poll it looks like there's a lot of interesting classification and data extraction it's great to see that you saw the value of classification and in data extraction today we have another question in the chat uh would we be able to read a PDF containing different invoices in different formats and be able to extract the information from the file and fill it somewhere else that that's going to depend how you've um kind of set up things so as I think you're referring to the fact that I said you can have a document with multiple Pages for classification so when you feed a document with multiple multiple pages to the document classifier you can classify is each page individually meaning that if you have a document that is should go to a specific extraction queue you can do that and now in your question it sounds like you might have different invoices in the same document and in that case it depends if you're actually processing them with different use cases or which is more likely you would have one single use case for all your invoices so in that case you'll still need to split the document between between the different invoices to to the extraction so I hope that answers the questions there's another question are all the fields mapped to a use case mandatory to have values after training for the document task studies to be changed to done so um all right there is the new feature that we talked about today which is required film so if you sell a field as required then you'll have to provide a value for the task really to be complete now if the value is missing from the document for some reason you can still flag that value as missing in the document that's something that's possible on the on the UI just mark it as it's missing in the document and in that case the the tasks are still shall be done asking if we can identify spelling errors or grammatical errors kind of like a proofreading for example that no that's not something we cover yet so we are focusing on uh classification and data extraction today so extracting data from documents so that you can uh use them in your workflows um so we don't have something for for proofreading all right and I think that's it for the questions thank you everybody for joining us uh again if you have questions don't hesitate to go to our community we'll be back in two weeks for another AI Academy session and in two weeks we'll talk about now assist for search so kicking off a series for our gen AI products thank you everybody for attending thank you Eliza for helping me manage the session today really appreciate it and I see you everybody in two weeks
https://www.youtube.com/watch?v=F4fTixbNmMY