Convert PDF to Word in .NET Smart Data Extractor
19 Sep 20267 minutes to read
Word (DOCX) is a widely used format for creating and editing professional documents. The Syncfusion® Smart Data Extractor library supports PDF and Images into Word conversion in .NET, enabling seamless transformation of PDF files into fully editable Word documents while preserving the original layout, barcodes, tables and images. This feature makes it easier to reuse content, improve accessibility, and integrate document data into .NET applications and business workflows.
Assemblies and NuGet packages required
Refer to the following links for the assemblies and NuGet packages required based on your target platform to convert PDF or Image as a Word file using the Syncfusion® Smart Data Extractor library.
Convert PDF or Image to Word Document
To convert a PDF document or Image into a Word using the ExtractDataAsWordDocument method of the DataExtractor class, refer to the following code example:
using Syncfusion.SmartDataExtractor;
using Syncfusion.DocIO.DLS;
//Open the input PDF file as a stream.
using (FileStream stream = new FileStream("Input.pdf", FileMode.Open, FileAccess.Read))
{
//Initialize the Data Extractor.
DataExtractor extractor = new DataExtractor();
//Extract data as WordDocument.
WordDocument word = extractor.ExtractDataAsWordDocument(stream);
//Save the extracted Word data into an output file.
word.Save("Output.docx");
word.Close();
}using Syncfusion.SmartDataExtractor;
using Syncfusion.DocIO.DLS;
//Open the input PDF file as a stream.
using (FileStream stream = new FileStream("Input.pdf", FileMode.Open, FileAccess.Read))
{
//Initialize the Data Extractor.
DataExtractor extractor = new DataExtractor();
//Extract data as WordDocument.
WordDocument word = extractor.ExtractDataAsWordDocument(stream);
//Save the output file.
word.Save("Output.docx");
word.Close();
}If you want to convert an image instead of a PDF, replace the input stream with the image file (for example, Input.jpg or Input.png). The rest of the code remains unchanged.
Convert a range of PDF pages to Word
To convert a specific range of PDF pages from the PDF document into Word by using the ExtractDataAsWordDocument method of the DataExtractor class, refer to the following code example:
using System.IO;
using Syncfusion.SmartDataExtractor;
using Syncfusion.DocIO.DLS;
//Open the input PDF file as a stream.
using (FileStream stream = new FileStream("Input.pdf", FileMode.Open, FileAccess.Read))
{
//Initialize the Data Extractor.
DataExtractor extractor = new DataExtractor();
//Set the page range for conversion (example: pages 2 to 4).
extractor.PageRange = new int[,] { { 2, 4 } };
//Convert the selected pages into a Word document.
WordDocument document = extractor.ExtractDataAsWordDocument(stream);
//Save the Word document.
document.Save("Output.docx");
document.Close();
}using System.IO;
using Syncfusion.SmartDataExtractor;
using Syncfusion.DocIO.DLS;
//Open the input PDF file as a stream.
using (FileStream stream = new FileStream("Input.pdf", FileMode.Open, FileAccess.Read))
{
//Initialize the Data Extractor.
DataExtractor extractor = new DataExtractor();
//Set the page range for conversion (example: pages 2 to 4).
extractor.PageRange = new int[,] { { 2, 4 } };
//Convert the selected pages to a Word document.
WordDocument document = extractor.ExtractDataAsWordDocument(stream);
//Save the Word document.
document.Save("Output.docx");
document.Close();
}If you want to convert an image instead of a PDF, replace the input stream with the image file (for example, Input.jpg or Input.png). The rest of the code remains unchanged.
Convert PDF or Image to HTML Document
To convert a PDF document or image into HTML output using the ExtractDataAsHtml method of the DataExtractor class, refer to the following code example:
using Syncfusion.SmartDataExtractor;
//Open the input PDF file as a stream.
using (FileStream stream = new FileStream("Input.pdf", FileMode.Open, FileAccess.Read))
{
//Initialize the Data Extractor.
DataExtractor extractor = new DataExtractor();
//Extract data as HTML.
string htmlContent = extractor.ExtractDataAsHtml(stream);
//Save the extracted data into the HTML file.
File.WriteAllText("Output.html", htmlContent);
}using Syncfusion.SmartDataExtractor;
//Open the input PDF file as a stream.
using (FileStream stream = new FileStream("Input.pdf", FileMode.Open, FileAccess.Read))
{
//Initialize the Data Extractor.
DataExtractor extractor = new DataExtractor();
//Extract data as HTML.
string htmlContent = extractor.ExtractDataAsHtml(stream);
//Save the extracted data into the HTML file.
File.WriteAllText("Output.html", htmlContent);
}If you want to convert an image instead of a PDF, replace the input stream with the image file (for example, Input.jpg or Input.png). The rest of the code remains unchanged.
Supported and Unsupported PDF Elements
The following table lists the PDF elements and their preservation details in the Word document.
| PDF Elements | Supported |
|---|---|
| Header, Paragraph Title, Document Title | Yes |
| Image | Yes |
| Table | Yes |
| Text Inline Styles | Yes (Bold and Italic) |
| Subscript, Superscript | No |
| Underline, Strikethrough | No |
| List | No (Converted as line-by-line text) |
| Charts and Barcodes | Yes (Preserved as images) |
| Code blocks, Footer, Page Number | Yes (Preserved as text) |
| Link | No (Preserved as plain text) |