dogear

enter for all results · esc to close

aws-pdf-textract-pipeline

github.comtool

ETL pipeline for crawling PDFs from the Web using Puppeteer and transforming their contents into structured data using AWS Textract and storing the results in DynamoDB.

from
CDK
added
2026-10-10
likes
0

CDK › Construct Libraries > Workflows: “ETL pipeline for crawling PDFs from the Web using Puppeteer and transforming their contents into structured data using AWS Textract and storing the results in DynamoDB.”