aws-pdf-textract-pipeline
ETL pipeline for crawling PDFs from the Web using Puppeteer and transforming their contents into structured data using AWS Textract and storing the results in DynamoDB.
- from
- CDK
- added
- 2026-10-10
- likes
- 0
CDK › Construct Libraries > Workflows: “ETL pipeline for crawling PDFs from the Web using Puppeteer and transforming their contents into structured data using AWS Textract and storing the results in DynamoDB.”