How to Convert a PDF to Text via Optical Character Recognition in Node.JS

Downloading an important PDF document only to have it not scan properly as text has been personal pet peeve both as a student and professionally. With the inability to copy and paste or edit the PDF, your time may be eaten up by tedious transcription that may have otherwise been avoided. The following API, however, will mitigate this issue using Optical Character Recognition to convert the PDF to text that can then be utilized more fully.

To start our function, we first need to run this command to install the SDK:

npm install cloudmersive-ocr-api-client --save

Or, you can add this snippet to your package.json:

"dependencies": {
"cloudmersive-ocr-api-client": "^1.3.3"

Then, we can call our function:

var CloudmersiveOcrApiClient = require('cloudmersive-ocr-api-client');
var defaultClient = CloudmersiveOcrApiClient.ApiClient.instance;

As you can see within the code block, the parameters for this API include the required image file (PDF file), as well as the optional recognition mode (basic, normal, and advanced), language, and preprocessing (image enhancement). The possible values for these parameters are included within the code as notes, so be sure that your inputs are valid according to these conditions.

The API Key for this function can be retrieved for free and with no commitment from the Cloudmersive website. This will give you access to 800 calls per month across our entire API library.

There’s an API for that. Cloudmersive is a leader in Highly Scalable Cloud APIs.