# Process Unstructured Content using AWS ML Services.

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690888012825/c56548e3-bfbf-43d4-b7b0-ccf2d4f26d26.png align="center")

In this project, We develop a robust model for processing unstructured content, extract valuable insights from the data, and enable automation and data-driven decision-making, thereby enhancing operational efficiency and facilitating informed decision-making within organizations.

AWS Services:

* Amazon S3
    
* Amazon Lambda
    
* Amazon Textract
    
* Amazon Comprehend
    

## Practical Implementation

We process resume documents from the Resume Entities for NER dataset to get insights such as candidates' skills by automating this workflow, we use Amazon Textract to extract text from these resumes and Amazon Comprehend custom entity recognition to detect skills such as AWS, C, and C++ as custom entities.

1. Launching CloudFormation Stack
    
    Step 1: Log in to AWS Console
    
    Step 2: [Cloudformation template](https://github.com/aws-samples/amazon-custom-entity-recognition-textract-comprehend/tree/master/cloud_formation) download this template in your local machine
    
    ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690886694769/fb7090dd-41ee-4ba6-8f60-a29fb1c1e236.png align="center")
    
    Step 3: Go to CloudFormation--&gt; Create Stack
    
    Step 4: Fill The details
    
    ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690886742111/1c09e7b1-e8d3-45b5-9eb0-7971fd057b4f.png align="center")
    
    Leave other things as it is.
    
    ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690886797766/227644d3-4204-430a-b7cb-0b06560084cc.png align="center")
    
    Wait for the stack to finish running.
    
    Step 5: On the Output tab of the Cloudformation stack, record the sagemaker URL.
    
    ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690886977369/894df6ec-fc63-4fdc-bc91-c850ca88bd4b.png align="center")
    
    ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690886984191/2e747f3c-883f-45f0-8704-88a13409c021.png align="center")
    
    1. Running Workflow on Jupyter Notebook
        
        Step 1: Open Sagemaker URL.
        
        Step 2: Under the New drop-down menu Choose Terminal
        
        Step 3: On the Terminal
        
        `cd sagemaker;`
        
        `clone` [`https://github.com/aws-samples/amazon-custom-entity-recognition-textract-comprehend`](https://github.com/aws-samples/amazon-custom-entity-recognition-textract-comprehend)
        
        ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690887234990/0d3a4c89-474f-4ba3-8bb8-df16c13ad73f.png align="center")
        
        Step 4: Open `Textract_Comprehend_Custom_Entity_Recognition.ipynb`
        
        Step 5: Run the Cells.
        
    2. After Running the cells. we have to check that Bucket is created or not
        
        ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690887348037/a8c90e83-2d84-4660-bce7-4515e64ba891.png align="center")
        
        ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690887360300/d8a1bd87-98f3-4ec4-91cc-2d1ff964d26a.png align="center")
        
        ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690887381648/8f4fc6df-8e12-47ea-8f45-1540d6b82994.png align="center")
        
        1. Go to Comprehend
            
            ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690887400188/5dcb4cae-febd-45f5-a4de-7b347b0b4ebc.png align="center")
            
            ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690887415394/79e10b1c-d9e5-47cf-80ef-ee3ca99bba29.png align="center")
            
            1. ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690887432320/44ae743b-1a08-470b-9b1a-b966e5a335b0.png align="center")
                
                ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690887452402/1617f4a8-3427-47ce-82e6-e23ca35bdef7.png align="center")
                
                ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690887479513/3e0320b1-a701-4225-880e-d8b9d2dd2e81.png align="center")
                
                ![](https://cdn.hashnode.com/res/hashnode/image/upload/v1690887489037/58e4be3d-803b-469f-9825-4d31f1dd370a.png align="center")
                
                Conclusion:
                
                The integration of Amazon Textract and Amazon Comprehend provides a powerful solution for extracting custom entities from documents. By leveraging Textract's document processing capabilities and Comprehend's entity recognition capabilities, businesses can automate the extraction of specific information that is relevant to their domain. This enables streamlined workflows, improved efficiency, and valuable insights from text-based documents.
