Convert a two-column research paper PDF to Word without mixing the left and right columns
How to Convert a Two-Column Research Paper PDF to Word in the Right Reading Order
Learn why text copied from a two-column PDF becomes scrambled and how to move headings, body text, footnotes, figures, and captions into Word in the correct reading order.
Visual layout and stored reading order are different things
When you read a two-column research paper, you naturally move down the left column and then continue at the top of the right. The characters inside the PDF, however, are not necessarily stored in that sequence.
A PDF can position text by page coordinates without encoding paragraphs or columns. The order produced by copy and paste depends on how the file was created, which is why you may see problems such as:
- Lines from the left and right columns alternate.
- A figure caption interrupts a paragraph.
- Page numbers and running headers appear as body text.
- A footnote appears before the sentence that references it.
- Words remain split by end-of-line hyphens.
Before exporting the paper to Word, check its reading order as well as whether the characters can be extracted.
Classify page layouts before conversion
A paper rarely uses one layout from beginning to end. Identify the main page types first:
| Page layout | Common example | What to watch for |
|---|---|---|
| Single column | Title and abstract | The page may switch to two columns below a full-width heading |
| Two columns | Main text | The boundary between the left and right columns must be detected |
| Two columns with a full-width figure | Results or discussion | Follow how the text continues before and after the figure |
| Footnotes | Bottom of a page | Treat them as a separate region from the body |
| References | End of the paper | Preserve one bibliographic entry at a time, not just visual lines |
Do not choose conversion settings from the title page alone. Inspect pages in the middle and at the end of the paper as well.
Reconstruct the correct reading order
1. Divide each page into regions
Separate full-width content, the left column, the right column, headers, footers, and footnotes. Splitting the page at its center line is not enough: a heading, equation, figure, or table spanning both columns would be cut in half.
Find full-width elements first, then process the two-column regions above and below them independently.
2. Order content within each column
Read elements from top to bottom in the left column, followed by the right column. Exclude page numbers and running heads from the body text.
Paragraph indentation, line spacing, and font size can help distinguish a heading or new paragraph. Sorting every object by coordinates alone may mix nearby figures and footnotes into the body.
3. Keep figures and captions together
When a figure or table spans both columns, the body text may stop before it and resume below it. Keep the caption with the figure instead of inserting it into the surrounding paragraph.
Even if you do not plan to reuse the image in Word, retain the caption and figure number. Otherwise, references such as “as shown in Figure 2” lose their target.
4. Separate footnotes from body text
Because footnotes sit at the bottom of a page, coordinate-based extraction may place them directly after the right-column text. In Word, create actual footnotes when possible, or at least label them clearly as a separate section.
Verify that every reference marker in the text points to the correct footnote number.
5. Repair line-break hyphenation carefully
Narrow columns often split English words with a hyphen at the end of a line. Removing every hyphen will also damage legitimately hyphenated terms. Consider removing one only when:
- The hyphen occurs at the visual end of a line.
- The next line begins with a lowercase letter.
- The combined form is used as one word in a dictionary or elsewhere in the paper.
Leave uncertain cases unchanged and flag them for review or targeted search and replace.
What to check after exporting to Word
Review layout boundaries before reading the entire converted document:
- The transition from abstract to main text
- The end of each left column and the top of its right column
- The end of one page and the start of the next
- Text immediately before and after a full-width figure or table
- Footnote markers and their corresponding notes
- Line breaks between reference entries
Checking these boundaries first exposes major reading-order errors quickly.
Should you export to Word or Markdown?
Choose Word when you need editable headings, paragraphs, tables, equations, and comments in reading order. Markdown is often more convenient for version comparisons or as input to another text-processing workflow.
Recreating the original two-column page and recovering the content in the correct order are separate goals. For editing, a single-column document with explicit reading order is usually easier to work with than a visual replica of the PDF.
PDFIntact’s free check reports the page count, text-layer condition, and detected tables and figures without uploading the PDF. It does not verify reading order. After conversion, inspect representative layouts and the boundary cases above before relying on the full document.
Check without uploading
See what can be extracted from this PDF first
The free check runs in your browser. Your original file is not uploaded.
Check your two-column PDF for free