How to Convert a PDF to Text via Optical Character Recognition in Node.JS

Downloading an important PDF document only to have it not scan properly as text has been personal pet peeve both as a student and professionally. With the inability to copy and paste or edit the PDF, your time may be eaten up by tedious transcription that may have otherwise been avoided. The following API, however, will mitigate this issue using Optical Character Recognition to convert the PDF to text that can then be utilized more fully.

Image for post
Image for post

To start our function, we first need to run this command to install the SDK:

Or, you can add this snippet to your package.json:

Then, we can call our function:

As you can see within the code block, the parameters for this API include the required image file (PDF file), as well as the optional recognition mode (basic, normal, and advanced), language, and preprocessing (image enhancement). The possible values for these parameters are included within the code as notes, so be sure that your inputs are valid according to these conditions.

The API Key for this function can be retrieved for free and with no commitment from the Cloudmersive website. This will give you access to 800 calls per month across our entire API library.

There’s an API for that. Cloudmersive is a leader in Highly Scalable Cloud APIs.

Get the Medium app

A button that says 'Download on the App Store', and if clicked it will lead you to the iOS App store
A button that says 'Get it on, Google Play', and if clicked it will lead you to the Google Play store