Extract text using TextLineCollection in ASP.NET Core
28 Feb 20261 minute to read
The Syncfusion PDF Viewer server Library extracts text and its precise bounding coordinates from a PDF page using the ExtractText method. The output TextLineCollection contains detailed layout information for each line of text.
Prerequisites
Ensure the following NuGet package is installed in the project:
- Syncfusion.EJ2.PdfViewer.AspNet.Core
NOTE
The Syncfusion.Pdf.Net.Core and Syncfusion.Compression.Net.Core packages are required dependencies and are typically installed automatically with the PDF Viewer package.
Implementation steps
Follow these steps to extract text and layout coordinates from a PDF:
- Load the PDF document into a
PdfLoadedDocumentinstance. - Access the desired page as a
PdfLoadedPageobject. - Call the
ExtractTextmethod, passing theTextLineCollectionas an output parameter.
var path = @"currentDirectory\..\..\..\..\Data\Simple.pdf";
var fileInfo = new FileInfo(path);
var docStream = new FileStream(fileInfo.FullName, FileMode.Open, FileAccess.Read);
// Load the PDF document.
PdfLoadedDocument document = new PdfLoadedDocument(docStream);
// Loading page collections
PdfPageBase page = document.Pages[0] as PdfLoadedPage;
//Extract text from the page.
var text = page.ExtractText(out TextLineCollection textLineCollection);Find the sample How to extract text using TextLineCollection.
NOTE
Ensure the document path and any output locations are valid for the hosting environment, and dispose of the
PdfLoadedDocumentafter extraction to release file handles.