{"id":164776,"date":"2019-11-23T16:01:54","date_gmt":"2019-11-23T10:31:54","guid":{"rendered":"https:\/\/www.digitalvidya.com\/blog\/?p=164776"},"modified":"2022-04-28T17:33:29","modified_gmt":"2022-04-28T12:03:29","slug":"work-with-python-pdf","status":"publish","type":"post","link":"https:\/\/www.digitalvidya.com\/blog\/work-with-python-pdf\/","title":{"rendered":"A Complete Guide on How to Work With a PDF in Python"},"content":{"rendered":"<p dir=\"ltr\">Python is a high-level language expressed with a simple syntax. This makes learning easy for new programmers. Some Python libraries can handle unstructured sources of data such as PDFs. Useful information such as audio, video, connections, buttons, business logic, and form fields can be found in PDFs.<\/p>\n<p dir=\"ltr\">For displaying and sharing files, PDF or Portable File Format is a file format. The PDF was developed by Adobe but is now maintained by the International Organization for Standardization (ISO). You must use the PyPDF2 package while dealing with Python&#8217;s PDF. It is a package of pure Python that can be used to perform various PDF operations.<\/p>\n<p dir=\"ltr\">Text analysis comes into play when a PDF is stored. Python is used to model a lot of the code and libraries for Text Analytics. Once the required information has been collected, the data can be used in the Natural Language Processing and Machine Learning system.<\/p>\n<div class=\"ast-oembed-container \" style=\"height: 100%;\"><iframe title=\"Building a PDF Data Extractor Using Python!!\" width=\"500\" height=\"281\" src=\"https:\/\/www.youtube.com\/embed\/UmPe07a3bWs?feature=oembed\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share\" referrerpolicy=\"strict-origin-when-cross-origin\" allowfullscreen><\/iframe><\/div>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">Here is a list of libraries that can be used for handling PDF files:<\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong>PDFMiner<\/strong> \u2013 This library is used to extract useful information from the PDF documents. Unlike other tools, the entire focus of this package is to get and analyze the data.<\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong>PyPDF2<\/strong> \u2013 This is a PDF library made of pure Python that can harvest, split, transform and merge PDFs together. There are also options available for adding custom data, passwords, and viewing options to PDF files. You can merge entire PDFs together and retrieve metadata and text from PDF.<\/span><\/p>\n<div class=\"cta-form\"><div class=\"content\"><div class=\"left-block\"><div class=\"section-title\">Want to Know the Path to Become a <strong> Data Science Expert?<\/strong><\/div><div class=\"desc\"><p>Download Detailed Brochure and Get Complimentary access to Live Online Demo Class with Industry Expert.<\/p>\n<\/div><div class=\"time\"> Date: August 1 (Sat) | 11 AM - 12 PM (IST)<\/div><\/div><div class=\"right-block\"><script>\nvar gform;gform||(document.addEventListener(\"gform_main_scripts_loaded\",function(){gform.scriptsLoaded=!0}),document.addEventListener(\"gform\/theme\/scripts_loaded\",function(){gform.themeScriptsLoaded=!0}),window.addEventListener(\"DOMContentLoaded\",function(){gform.domLoaded=!0}),gform={domLoaded:!1,scriptsLoaded:!1,themeScriptsLoaded:!1,isFormEditor:()=>\"function\"==typeof InitializeEditor,callIfLoaded:function(o){return!(!gform.domLoaded||!gform.scriptsLoaded||!gform.themeScriptsLoaded&&!gform.isFormEditor()||(gform.isFormEditor()&&console.warn(\"The use of gform.initializeOnLoaded() is deprecated in the form editor context and will be removed in Gravity Forms 3.1.\"),o(),0))},initializeOnLoaded:function(o){gform.callIfLoaded(o)||(document.addEventListener(\"gform_main_scripts_loaded\",()=>{gform.scriptsLoaded=!0,gform.callIfLoaded(o)}),document.addEventListener(\"gform\/theme\/scripts_loaded\",()=>{gform.themeScriptsLoaded=!0,gform.callIfLoaded(o)}),window.addEventListener(\"DOMContentLoaded\",()=>{gform.domLoaded=!0,gform.callIfLoaded(o)}))},hooks:{action:{},filter:{}},addAction:function(o,r,e,t){gform.addHook(\"action\",o,r,e,t)},addFilter:function(o,r,e,t){gform.addHook(\"filter\",o,r,e,t)},doAction:function(o){gform.doHook(\"action\",o,arguments)},applyFilters:function(o){return gform.doHook(\"filter\",o,arguments)},removeAction:function(o,r){gform.removeHook(\"action\",o,r)},removeFilter:function(o,r,e){gform.removeHook(\"filter\",o,r,e)},addHook:function(o,r,e,t,n){null==gform.hooks[o][r]&&(gform.hooks[o][r]=[]);var d=gform.hooks[o][r];null==n&&(n=r+\"_\"+d.length),gform.hooks[o][r].push({tag:n,callable:e,priority:t=null==t?10:t})},doHook:function(r,o,e){var t;if(e=Array.prototype.slice.call(e,1),null!=gform.hooks[r][o]&&((o=gform.hooks[r][o]).sort(function(o,r){return o.priority-r.priority}),o.forEach(function(o){\"function\"!=typeof(t=o.callable)&&(t=window[t]),\"action\"==r?t.apply(null,e):e[0]=t.apply(null,e)})),\"filter\"==r)return e[0]},removeHook:function(o,r,t,n){var e;null!=gform.hooks[o][r]&&(e=(e=gform.hooks[o][r]).filter(function(o,r,e){return!!(null!=n&&n!=o.tag||null!=t&&t!=o.priority)}),gform.hooks[o][r]=e)}});\n<\/script>\n\n                <div class='gf_browser_unknown gform_wrapper gravity-theme gform-theme--no-framework dv-form_wrapper' data-form-theme='gravity-theme' data-form-index='0' id='gform_wrapper_287' ><div id='gf_287' class='gform_anchor' tabindex='-1'><\/div><form method='post' enctype='multipart\/form-data' target='gform_ajax_frame_287' id='gform_287' class='dv-form' action='\/blog\/wp-json\/wp\/v2\/posts\/164776#gf_287' data-formid='287' novalidate>\n                        <div class='gform-body gform_body'><div id='gform_fields_287' class='gform_fields top_label form_sublabel_below description_below validation_below'><div id=\"field_287_28\" class=\"gfield gfield--type-text gfield_contains_required field_sublabel_below gfield--no-description field_description_below hidden_label field_validation_below gfield_visibility_visible\"  ><label class='gfield_label gform-field-label' for='input_287_28'>Name<span class=\"gfield_required\"><span class=\"gfield_required gfield_required_text\">(Required)<\/span><\/span><\/label><div class='ginput_container ginput_container_text'><input name='input_28' id='input_287_28' type='text' value='' class='large'    placeholder='Name *' aria-required=\"true\" aria-invalid=\"false\"   \/><\/div><\/div><div id=\"field_287_2\" class=\"gfield gfield--type-email gfield_contains_required field_sublabel_below gfield--no-description field_description_below hidden_label field_validation_below gfield_visibility_visible\"  ><label class='gfield_label gform-field-label' for='input_287_2'>Email<span class=\"gfield_required\"><span class=\"gfield_required gfield_required_text\">(Required)<\/span><\/span><\/label><div class='ginput_container ginput_container_email'>\n                            <input name='input_2' id='input_287_2' type='email' value='' class='large'   placeholder='Email *' aria-required=\"true\" aria-invalid=\"false\"  \/>\n                        <\/div><\/div><div id=\"field_287_3\" class=\"gfield gfield--type-phone gfield_contains_required field_sublabel_below gfield--no-description field_description_below hidden_label field_validation_below gfield_visibility_visible\"  ><label class='gfield_label gform-field-label' for='input_287_3'>Phone<span class=\"gfield_required\"><span class=\"gfield_required gfield_required_text\">(Required)<\/span><\/span><\/label><div class='ginput_container ginput_container_phone'><input name='input_3' id='input_287_3' type='tel' value='' class='large'  placeholder='Phone *' aria-required=\"true\" aria-invalid=\"false\"   \/><\/div><\/div><div id=\"field_287_34\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_34' id='input_287_34' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='fmc' \/><\/div><\/div><div id=\"field_287_9\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_9' id='input_287_9' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='' \/><\/div><\/div><div id=\"field_287_10\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_10' id='input_287_10' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='' \/><\/div><\/div><div id=\"field_287_11\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_11' id='input_287_11' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='' \/><\/div><\/div><div id=\"field_287_12\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_12' id='input_287_12' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='https:\/\/www.digitalvidya.com\/blog\/wp-json\/wp\/v2\/posts\/164776' \/><\/div><\/div><div id=\"field_287_13\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_13' id='input_287_13' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='' \/><\/div><\/div><div id=\"field_287_14\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_14' id='input_287_14' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='' \/><\/div><\/div><div id=\"field_287_15\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_15' id='input_287_15' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='' \/><\/div><\/div><div id=\"field_287_17\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_17' id='input_287_17' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='' \/><\/div><\/div><div id=\"field_287_18\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_18' id='input_287_18' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='' \/><\/div><\/div><div id=\"field_287_19\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_19' id='input_287_19' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='' \/><\/div><\/div><div id=\"field_287_20\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_20' id='input_287_20' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='{dmo_day}' \/><\/div><\/div><div id=\"field_287_25\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_25' id='input_287_25' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='{dmo_date}' \/><\/div><\/div><div id=\"field_287_24\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_24' id='input_287_24' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='{dmo_time}' \/><\/div><\/div><div id=\"field_287_35\" class=\"gfield gfield--type-hidden gfield--width-full gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_35' id='input_287_35' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='{dmo_time}' \/><\/div><\/div><div id=\"field_287_36\" class=\"gfield gfield--type-hidden gfield--width-full gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_36' id='input_287_36' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='' \/><\/div><\/div><div id=\"field_287_37\" class=\"gfield gfield--type-hidden gfield--width-full gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_37' id='input_287_37' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='' \/><\/div><\/div><div id=\"field_287_22\" class=\"gfield gfield--type-hidden gform_hidden field_sublabel_below gfield--no-description field_description_below field_validation_below gfield_visibility_visible\"  ><div class='ginput_container ginput_container_text'><input name='input_22' id='input_287_22' type='hidden' class='gform_hidden'  aria-invalid=\"false\" value='' \/><\/div><\/div><div id=\"field_287_38\" class=\"gfield gfield--type-honeypot gform_validation_container field_sublabel_below gfield--has-description field_description_below field_validation_below gfield_visibility_visible\"  ><label class='gfield_label gform-field-label' for='input_287_38'>Email<\/label><div class='ginput_container'><input name='input_38' id='input_287_38' type='text' value='' autocomplete='new-password'\/><\/div><div class='gfield_description' id='gfield_description_287_38'>This field is for validation purposes and should be left unchanged.<\/div><\/div><\/div><\/div>\n        <div class='gform-footer gform_footer top_label'> <input type='submit' id='gform_submit_button_287' class='gform_button button' onclick='gform.submission.handleButtonClick(this);' data-submission-type='submit' value='Register Now'  \/> <input type='hidden' name='gform_ajax' value='form_id=287&amp;title=&amp;description=&amp;tabindex=0&amp;theme=gravity-theme&amp;styles=[]&amp;hash=0c9c234943f312d5985375593e68ff15' \/>\n            <input type='hidden' class='gform_hidden' name='gform_submission_method' data-js='gform_submission_method_287' value='iframe' \/>\n            <input type='hidden' class='gform_hidden' name='gform_theme' data-js='gform_theme_287' id='gform_theme_287' value='gravity-theme' \/>\n            <input type='hidden' class='gform_hidden' name='gform_style_settings' data-js='gform_style_settings_287' id='gform_style_settings_287' value='[]' \/>\n            <input type='hidden' class='gform_hidden' name='is_submit_287' value='1' \/>\n            <input type='hidden' class='gform_hidden' name='gform_submit' value='287' \/>\n            \n            <input type='hidden' class='gform_hidden' name='gform_unique_id' value='' \/>\n            <input type='hidden' class='gform_hidden' name='state_287' value='WyJbXSIsIjc5YTc5ZmI0NzJmZjc1YWM4NDI0ZmE2ZmFkNDQwNjAxIl0=' \/>\n            <input type='hidden' autocomplete='off' class='gform_hidden' name='gform_target_page_number_287' id='gform_target_page_number_287' value='0' \/>\n            <input type='hidden' autocomplete='off' class='gform_hidden' name='gform_source_page_number_287' id='gform_source_page_number_287' value='1' \/>\n            <input type='hidden' name='gform_field_values' value='' \/>\n            \n        <\/div>\n                        <p style=\"display: none !important;\" class=\"akismet-fields-container\" data-prefix=\"ak_\"><label>&#916;<textarea name=\"ak_hp_textarea\" cols=\"45\" rows=\"8\" maxlength=\"100\"><\/textarea><\/label><input type=\"hidden\" id=\"ak_js_1\" name=\"ak_js\" value=\"87\"\/><script>\ndocument.getElementById( \"ak_js_1\" ).setAttribute( \"value\", ( new Date() ).getTime() );\n<\/script>\n<\/p><\/form>\n                        <\/div>\n\t\t                <iframe style='display:none;width:0px;height:0px;' src='about:blank' name='gform_ajax_frame_287' id='gform_ajax_frame_287' title='This iframe contains the logic required to handle Ajax powered Gravity Forms.'><\/iframe>\n\t\t                <script>\ngform.initializeOnLoaded( function() {gformInitSpinner( 287, 'https:\/\/www.digitalvidya.com\/blog\/wp-content\/plugins\/gravityforms\/images\/spinner.svg', true );jQuery('#gform_ajax_frame_287').on('load',function(){var contents = jQuery(this).contents().find('*').html();var is_postback = contents.indexOf('GF_AJAX_POSTBACK') >= 0;if(!is_postback){return;}var form_content = jQuery(this).contents().find('#gform_wrapper_287');var is_confirmation = jQuery(this).contents().find('#gform_confirmation_wrapper_287').length > 0;var is_redirect = contents.indexOf('gformRedirect(){') >= 0;var is_form = form_content.length > 0 && ! is_redirect && ! is_confirmation;var mt = parseInt(jQuery('html').css('margin-top'), 10) + parseInt(jQuery('body').css('margin-top'), 10) + 100;if(is_form){jQuery('#gform_wrapper_287').html(form_content.html());if(form_content.hasClass('gform_validation_error')){jQuery('#gform_wrapper_287').addClass('gform_validation_error');} else {jQuery('#gform_wrapper_287').removeClass('gform_validation_error');}setTimeout( function() { \/* delay the scroll by 50 milliseconds to fix a bug in chrome *\/ jQuery(document).scrollTop(jQuery('#gform_wrapper_287').offset().top - mt); }, 50 );if(window['gformInitDatepicker']) {gformInitDatepicker();}if(window['gformInitPriceFields']) {gformInitPriceFields();}var current_page = jQuery('#gform_source_page_number_287').val();gformInitSpinner( 287, 'https:\/\/www.digitalvidya.com\/blog\/wp-content\/plugins\/gravityforms\/images\/spinner.svg', true );jQuery(document).trigger('gform_page_loaded', [287, current_page]);window['gf_submitting_287'] = false;}else if(!is_redirect){var confirmation_content = jQuery(this).contents().find('.GF_AJAX_POSTBACK').html();if(!confirmation_content){confirmation_content = contents;}jQuery('#gform_wrapper_287').replaceWith(confirmation_content);jQuery(document).scrollTop(jQuery('#gf_287').offset().top - mt);jQuery(document).trigger('gform_confirmation_loaded', [287]);window['gf_submitting_287'] = false;wp.a11y.speak(jQuery('#gform_confirmation_message_287').text());}else{jQuery('#gform_287').append(contents);if(window['gformRedirect']) {gformRedirect();}}jQuery(document).trigger(\"gform_pre_post_render\", [{ formId: \"287\", currentPage: \"current_page\", abort: function() { this.preventDefault(); } }]);        if (event && event.defaultPrevented) {                return;        }        const gformWrapperDiv = document.getElementById( \"gform_wrapper_287\" );        if ( gformWrapperDiv ) {            const visibilitySpan = document.createElement( \"span\" );            visibilitySpan.id = \"gform_visibility_test_287\";            gformWrapperDiv.insertAdjacentElement( \"afterend\", visibilitySpan );        }        const visibilityTestDiv = document.getElementById( \"gform_visibility_test_287\" );        let postRenderFired = false;        function triggerPostRender() {            if ( postRenderFired ) {                return;            }            postRenderFired = true;            gform.core.triggerPostRenderEvents( 287, current_page );            if ( visibilityTestDiv ) {                visibilityTestDiv.parentNode.removeChild( visibilityTestDiv );            }        }        function debounce( func, wait, immediate ) {            var timeout;            return function() {                var context = this, args = arguments;                var later = function() {                    timeout = null;                    if ( !immediate ) func.apply( context, args );                };                var callNow = immediate && !timeout;                clearTimeout( timeout );                timeout = setTimeout( later, wait );                if ( callNow ) func.apply( context, args );            };        }        const debouncedTriggerPostRender = debounce( function() {            triggerPostRender();        }, 200 );        if ( visibilityTestDiv && visibilityTestDiv.offsetParent === null ) {            const observer = new MutationObserver( ( mutations ) => {                mutations.forEach( ( mutation ) => {                    if ( mutation.type === 'attributes' && visibilityTestDiv.offsetParent !== null ) {                        debouncedTriggerPostRender();                        observer.disconnect();                    }                });            });            observer.observe( document.body, {                attributes: true,                childList: false,                subtree: true,                attributeFilter: [ 'style', 'class' ],            });        } else {            triggerPostRender();        }    } );} );\n<\/script>\n<\/div><\/div><\/div>\n<p><span style=\"font-weight: 400\"><strong>Tabula-py<\/strong> \u2013 It is the tabula-java\u2019s Python wrapper which can be used for reading the tables present in PDF. You can also convert them into DataFrame of Pandas. There is also an option for converting the PDF file into JSON\/TSV\/CSV file.<\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong>Slate<\/strong> \u2013 It is PDFMiner\u2019s wrapper implementation.<\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong>PDFQuery<\/strong> \u2013 It is the light wrapper around pyquery, lxml, and pdfminer. With this, you can extract the data from PDFs reliable without writing long codes.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong>Xpdf<\/strong> \u2013 It is the Python wrapper that is currently offering just the utility to convert pdf to text.<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">The first pyPDF package was released in 2005. The last update to that package was made in 2010. Then, a company named Phasit created a package named PyPDF2 as a fork of pyPDF. This package was backwards compatible with pyPDF and worked perfectly for several years up to 2016. Then there were a few releases of pyPDF3 which was renamed to PyPDF4 later on. <\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">Almost all of these packages do at the same time. However, there is one major difference between PyPDF2+ and the original pyPDF which is that the former supports Python 3. Even though PyPDF2 was abandoned recently, PyPDF4 is not backwards compatible with it\u00a0<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">An alternative to PyPDF2 was created by Patrick Maupin with the name pdfrw. It does most of the things that PyPDF does. The only major difference between the two is that with pdfrw, you can integrate it with ReportLab package that can create a new PDF on ReportLab containing some or all part of a preexisting PDF.<\/span><\/p>\n<p dir=\"ltr\">\n<p dir=\"ltr\"><span style=\"font-weight: 400\">The first step for working with a PDF in Python is installing the package. You can use conda (if you are using Anaconda) or pip (if you are using regular Python) for installing PyPDF2. Here is what you need to do for installing PyPDF2 using pip:<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">$ pip install pypdf2<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">The installation process does not take much time as the PyPDF2 package doesn\u2019t have any dependencies. Now, let\u2019s move on to extracting information from PDF.<\/span><\/p>\n<h2><span class=\"ez-toc-section\" id=\"extracting\"><\/span><b>Extracting\u00a0<\/b><span class=\"ez-toc-section-end\"><\/span><\/h2>\n<figure id=\"attachment_164802\" aria-describedby=\"caption-attachment-164802\" style=\"width: 497px\" class=\"wp-caption aligncenter\"><img decoding=\"async\" src=\"https:\/\/www.digitalvidya.com\/blog\/wp-content\/uploads\/2019\/09\/PDF-to-Python_c0b0ea09308397eed080d56bb257dc90.png\" alt=\"Extraction Text from PDF\" class=\"wp-image-164802\" width=\"497\" height=\"285\" title=\"\"><figcaption id=\"caption-attachment-164802\" class=\"wp-caption-text\">Extraction text from pdf source &#8211; pdf tables<\/figcaption><\/figure>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">With the PyPDF2, you will be able to extract text and metadata from PDF. This comes in handy when you are working on automating the preexisting PDF files. You can extract the following types of data using the PyPDF2 package:<\/span><\/p>\n<p style=\"padding-left: 40px\"><span style=\"font-weight: 400\">\u21d2 Creator<\/span><\/p>\n<p style=\"padding-left: 40px\"><span style=\"font-weight: 400\">\u21d2 Author<\/span><\/p>\n<p style=\"padding-left: 40px\"><span style=\"font-weight: 400\">\u21d2 Subject<\/span><\/p>\n<p style=\"padding-left: 40px\"><span style=\"font-weight: 400\">\u21d2 Producer<\/span><\/p>\n<p style=\"padding-left: 40px\"><span style=\"font-weight: 400\">\u21d2 Title<\/span><\/p>\n<p style=\"padding-left: 40px\"><span style=\"font-weight: 400\">\u21d2 Number of Pages<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">To practice this, you need to get a PDF. Any PDF will do the job. In this example, let\u2019s assume that the name of the pdf is example.pdf. Now, here is the code that will get you access to the attributes of the PDF:<\/span><\/p>\n<pre><span style=\"font-weight: 400\"># extract_doc_info.py<\/span>\r\n\r\n<span style=\"font-weight: 400\">from PyPDF2 import PdfFileReader<\/span>\r\n\r\n<span style=\"font-weight: 400\">def extract_information(pdf_path):<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0with open(pdf_path, 'rb') as f:<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0pdf = PdfFileReader(f)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0information = pdf.getDocumentInfo()<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0number_of_pages = pdf.getNumPages()<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0txt = f\"\"\"<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0Information about {pdf_path}:\u00a0<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0Author: {information.author}<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0Creator: {information.creator}<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0Producer: {information.producer}<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0Subject: {information.subject}<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0Title: {information.title}<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0Number of pages: {number_of_pages}<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\"\"\"<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0print(txt)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0return information<\/span>\r\n\r\n<span style=\"font-weight: 400\">if __name__ == '__main__':<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0path = 'example.pdf'<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0extract_information(path)<\/span>\r\n\r\n<span style=\"font-weight: 400\">Here, you have used the PyPDF2 package for importing PdfFileReader. It is a class containing different methods used to interact with PDF files. In the above example, the instance of DocumentInformation is returned after calling .getDocumentInfo(). <\/span><\/pre>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">All the information you need on the PDF can be extracted by this. For returning the number of pages, you need to call .getNumPages().<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">The information variable used in the above example has attributes that can be used for extracting the remaining metadata from the document. You can even print the information and save it for future use.<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">There is an .extractText() function present in the PyPDF package that can be used for extracting text on the page objects.<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\"> However, many times this method turns out to be unsuccessful. In some PDF, you will get the text and in other cases, you will get an empty string. The best package for extracting text from PDF in Python is the PDFMiner project which is more robust and is designed specifically to extract from PDF.<\/span><\/p>\n<h2><span class=\"ez-toc-section\" id=\"rotating\"><\/span><b>Rotating<\/b><span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">More than often you would have to deal with PDFs whose pages are in landscape mode instead of portrait mode. Then can even be upside down. This happens when someone creates a document by scanning them. With Python, you will be able to rotate these pages. <\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">Here is an example through which you will be able to understand how to rotate a few pages of a PDF with the PyPDF2 package:<\/span><\/p>\n<pre><span style=\"font-weight: 400\"># rotate_pages.py<\/span>\r\n\r\n<span style=\"font-weight: 400\">from PyPDF2 import PdfFileReader, PdfFileWriter<\/span>\r\n\r\n<span style=\"font-weight: 400\">def rotate_pages(pdf_path):<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0pdf_writer = PdfFileWriter()<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0pdf_reader = PdfFileReader(path)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0# Rotate page 90 degrees to the right<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0page_1 = pdf_reader.getPage(0).rotateClockwise(90)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0pdf_writer.addPage(page_1)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0# Rotate page 90 degrees to the left<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0page_2 = pdf_reader.getPage(1).rotateCounterClockwise(90)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0pdf_writer.addPage(page_2)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0# Add a page in normal orientation<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0pdf_writer.addPage(pdf_reader.getPage(2))<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0with open('rotate_pages.pdf', 'wb') as fh:<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0pdf_writer.write(fh)<\/span>\r\n\r\n<span style=\"font-weight: 400\">if __name__ == '__main__':<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0path = 'example.pdf'<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0rotate_pages(path)<\/span><\/pre>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">In this case, apart from the PdfFileReader, you will also have to import PdfFileWriter as you will have to write a new PDF. The pages that you want to modify are taken in the path through rotate_pages(). This also requires creating a writer object named pdf_writer and a reader object named pdf_reader within the function. Next, you have to get the desired pages for modifications through .<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">GetPage(). In the above example, we have started from the first page, which is page zero. Then, you pass in 90 degrees after calling .rotateClockwise(), the page\u2019s object. For page two you pass 90 degrees as well after calling .rotateCounterClockwise(). With PyPDF2, you can rotate a page only in increments of 90 degrees. Any other thing would raise an AssertionError.<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">After every call that you make to the rotation method, you need to call .addPage(). This is done for adding the page\u2019s rotated version to the writer object. The last step is using the .write() for writing out the new PDF. The parameter in this function is a file-like object.\u00a0<\/span><\/p>\n<h2><span class=\"ez-toc-section\" id=\"merging\"><\/span><b>Merging<\/b><span class=\"ez-toc-section-end\"><\/span><\/h2>\n<figure id=\"attachment_164803\" aria-describedby=\"caption-attachment-164803\" style=\"width: 516px\" class=\"wp-caption aligncenter\"><img decoding=\"async\" src=\"https:\/\/www.digitalvidya.com\/blog\/wp-content\/uploads\/2019\/09\/Merging-PDF_0269c00d99bb2e15d7ea1c9884d9cfac.jpg\" alt=\"Merging PDF\" class=\"wp-image-164803\" width=\"516\" height=\"315\" title=\"\"><figcaption id=\"caption-attachment-164803\" class=\"wp-caption-text\">Merging pdf source &#8211; drupal<\/figcaption><\/figure>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">With the PyPDF2 package, you will be able to merge two or more PDFs into a single PDF document. For example, you have several types of reports that need to have a standard cover page. To deal with this type of situation, you might need the help of Python and the PyPDF2 package.<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">Here, we have mentioned an example where you will be merging PDFs together.<\/span><\/p>\n<pre><span style=\"font-weight: 400\"># pdf_merging.py<\/span>\r\n\r\n<span style=\"font-weight: 400\">from PyPDF2 import PdfFileReader, PdfFileWriter<\/span>\r\n\r\n<span style=\"font-weight: 400\">def merge_pdfs(paths, output):<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0pdf_writer = PdfFileWriter()<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0for path in paths:<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0pdf_reader = PdfFileReader(path)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0for page in range(pdf_reader.getNumPages()):<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0# Add each page to the writer object<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0pdf_writer.addPage(pdf_reader.getPage(page))<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0# Write out the merged PDF<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0with open(output, 'wb') as out:<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0pdf_writer.write(out)<\/span>\r\n\r\n<span style=\"font-weight: 400\">if __name__ == '__main__':<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0paths = ['document1.pdf', 'document2.pdf']<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0merge_pdfs(paths, output='merged.pdf')<\/span><\/pre>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">The merge_pdfs() is used when you want to merge a list of PDFs together. You must be aware of the location where you want to save the result. This function takes the input path\u2019s list and the output for it to save the merged output. <\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">As you can see, a loop is created for the inputs and a PDF reader object is created for every input. The next step is iterating over the pages of the PDF file and add all the pages to itself using the .addPage(). After all the pages have been iterated of all the PDFs, the end result is written onto a single PDF.\u00a0<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">Another feature of PyPDF2 is that if you don\u2019t want to merge all the pages of the PDF and want to add just a range of pages, you can enhance the script. You can also use the argparse module or Python for creating a command-line interface for the function.<\/span><\/p>\n<h2><span class=\"ez-toc-section\" id=\"splitting\"><\/span><b>Splitting<\/b><span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">The opposite of merging, splitting is taking out a couple of pages from a PDF document. This is very beneficial when you are working with PDFs that have a lot of scanned-in content that might be repeated, you might not just need it or any other good reason that you might have to split the PDF file. <\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">Here is an example of splitting a single PDF into multiple files using PyPDF2:<\/span><\/p>\n<pre><span style=\"font-weight: 400\"># pdf_splitting.py<\/span>\r\n\r\n<span style=\"font-weight: 400\">from PyPDF2 import PdfFileReader, PdfFileWriter<\/span>\r\n\r\n<span style=\"font-weight: 400\">def split(path, name_of_split):<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0pdf = PdfFileReader(path)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0for page in range(pdf.getNumPages()):<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0pdf_writer = PdfFileWriter()<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0pdf_writer.addPage(pdf.getPage(page))<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0output = f'{name_of_split}{page}.pdf'<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0with open(output, 'wb') as output_pdf:<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0pdf_writer.write(output_pdf)<\/span>\r\n\r\n<span style=\"font-weight: 400\">if __name__ == '__main__':<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0path = 'Jupyter_Notebook_An_Introduction.pdf'<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0split(path, 'jupyter_page')<\/span><\/pre>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">As you can see in the above example, a PDF reader object is created and then a loop for all the pages. A new PDF writer instance is created and a single page is added for every page of the PDF. Then, a uniquely named file is used for writing the page out. After the script is done running, you will have every page of the PDF split into multiple PDFs.\u00a0<\/span><\/p>\n<h2><span class=\"ez-toc-section\" id=\"adding-a-watermark\"><\/span><b>Adding a Watermark<\/b><span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">Watermarks are a way to identify patterns and images on digital and printed documents. There are some watermarks that can be seen in just special lighting conditions. Watermarks are an overlay that is really important as they allow protection of intellectual properties like your PDFs or images. <\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">For watermarking your documents you can take the help of Python and the PyPDF2 package. To practice this, you need to have a watermark text or an image to use on the PDF. Take a look at this example:<\/span><\/p>\n<pre><span style=\"font-weight: 400\"># pdf_watermarker.py<\/span>\r\n\r\n<span style=\"font-weight: 400\">from PyPDF2 import PdfFileWriter, PdfFileReader<\/span>\r\n\r\n<span style=\"font-weight: 400\">def create_watermark(input_pdf, output, watermark):<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0watermark_obj = PdfFileReader(watermark)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0watermark_page = watermark_obj.getPage(0)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0pdf_reader = PdfFileReader(input_pdf)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0pdf_writer = PdfFileWriter()<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0# Watermark all the pages<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0for page in range(pdf_reader.getNumPages()):<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0page = pdf_reader.getPage(page)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0page.mergePage(watermark_page)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0pdf_writer.addPage(page)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0with open(output, 'wb') as out:<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0pdf_writer.write(out)<\/span>\r\n\r\n<span style=\"font-weight: 400\">if __name__ == '__main__':<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0create_watermark(<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0input_pdf='Jupyter_Notebook_An_Introduction.pdf',\u00a0<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0output='watermarked_notebook.pdf',<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0watermark='watermark.pdf')<\/span><\/pre>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">There are three arguments that can be accepted by create_watermark():<\/span><\/p>\n<p><b>Input_pdf<\/b><span style=\"font-weight: 400\">: This is the PDF file on which you have to put the watermark.<\/span><\/p>\n<p><b>Output_pdf<\/b><span style=\"font-weight: 400\">: This is the path where you will save the PDF with the watermark.<\/span><\/p>\n<p><b>Watermark<\/b><span style=\"font-weight: 400\">: This is the PDF where you have saved your watermark text or image.<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">As you can see in the code, you have to open the watermark PDF and take the first page of the document where the watermark is present. The next step is creating a PDF reader object using an input_pdf and a pdr-writer object to write the PDF with the watermark.<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\"> After this, you have to iterate all the pages in the input_pdf. You pass the watermark_page after calling the .mergePage(). This will place the watermark_page on the current page. The last step is to use the pdf_writer object for adding the newly merged page to the PDF and voila! You will have your PDF with the watermark.<\/span><\/p>\n<p dir=\"ltr\">\n<h2><span class=\"ez-toc-section\" id=\"encryption\"><\/span><b>Encryption<\/b><span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">Currently, you can just add a user and an owner password using the PyPDF2 package. With the owner password, you will have administrative privileges on the PDF. You will also be able to set permissions on the document. The user password allows you to just read the document.<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">With the PyPDF2, you can set the owner password even though you can set any permission on the document. So, for encrypting the PDF, you can just add the password. Take a look at this example:<\/span><\/p>\n<pre><span style=\"font-weight: 400\"># pdf_encrypt.py<\/span>\r\n\r\n<span style=\"font-weight: 400\">from PyPDF2 import PdfFileWriter, PdfFileReader<\/span>\r\n\r\n<span style=\"font-weight: 400\">def add_encryption(input_pdf, output_pdf, password):<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0pdf_writer = PdfFileWriter()<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0pdf_reader = PdfFileReader(input_pdf)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0for page in range(pdf_reader.getNumPages()):<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0pdf_writer.addPage(pdf_reader.getPage(page))<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0pdf_writer.encrypt(user_pwd=password, owner_pwd=None,\u00a0<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0use_128bit=True)<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0with open(output_pdf, 'wb') as fh:<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0pdf_writer.write(fh)<\/span>\r\n\r\n<span style=\"font-weight: 400\">if __name__ == '__main__':<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0add_encryption(input_pdf='reportlab-sample.pdf',<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0output_pdf='reportlab-encrypted.pdf',<\/span>\r\n\r\n<span style=\"font-weight: 400\">\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0password='twofish')<\/span><\/pre>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">The add_encryption() uses the PDF paths for input as well as output and also the password that you have to add to the PDF. Next, a PDF writer is opened and then a reader object. Now, you will have to take an iteration of all the pages of the PDF to create a loop and add them to the writer for encrypting the complete input PDF. <\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">The last step is calling the.encrypt() where you have to put in the owner password, the user password, and whether you want the 128-bit encryption for the PDF file or not. The default setting is the 128-encryption turned on. You will have to set it to False for setting the 40-bit encryption. <\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">According to pdflib.com, the encryption used in PDF is either AES (Advanced Encryption Standard) or RC4. But you have to remember that even after encrypting your PDF, it doesn\u2019t mean that it is secure. There are several tools available that can remove passwords.\u00a0<\/span><\/p>\n<h2><span class=\"ez-toc-section\" id=\"reading-table-data\"><\/span><b>Reading Table Data<\/b><span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">For reading table data, you have to use the Tabula-py. The first step is installing it first through the following command:<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">pip install tabula-py<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">Here is what you need to do is extract the data:<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">import tabula<\/span><\/p>\n<pre><span style=\"font-weight: 400\"># reading the PDF file that contains Table Data<\/span><span style=\"font-weight: 400\">\r\n<\/span><span style=\"font-weight: 400\"># you can find find the pdf file with complete code in below<\/span><span style=\"font-weight: 400\">\r\n<\/span><span style=\"font-weight: 400\"># read_pdf will save the pdf table into Pandas Dataframe<\/span>\r\n\r\n<span style=\"font-weight: 400\">df = tabula.read_pdf(\"offense.pdf\")<\/span>\r\n\r\n<span style=\"font-weight: 400\"># in order to print first 5 lines of Table<\/span>\r\n\r\n<span style=\"font-weight: 400\">df.head()<\/span><\/pre>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">If there are multiple files present in the PDF file, you have to use the following command:<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">df = tabula.read_pdf(\u201coffense.pdf\u201d,multiple_tables=True)<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">For extracting specific information from a specific page of the PDF, you need to use this:<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">tabula.read_pdf(&#8220;offense.pdf&#8221;, area=(126,149,212,462), pages=1)<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">For putting the output into a JSON format, you need to try this:<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">tabula.read_pdf(&#8220;offense.pdf&#8221;, output_format=&#8221;json&#8221;)<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">Use the following command for converting the PDF into a CSV or an Excel file:<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">tabula.convert_into(&#8220;offense.pdf&#8221;, &#8220;offense_testing.xlsx&#8221;, output_format=&#8221;xlsx&#8221;)<\/span><\/p>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">To understand more about working with PDF packages, you can try the following resources:<\/span><\/p>\n<ul>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">The Github page for\u00a0<\/span><span style=\"font-weight: 400\">PyPDF4<\/span><\/li>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">The ReportLab website<\/span><\/li>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">The\u00a0<\/span><span style=\"font-weight: 400\">PyPDF2<\/span><span style=\"font-weight: 400\">\u00a0website<\/span><\/li>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">Camelot: PDF Table Extraction for Humans<\/span><\/li>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">The Github page for\u00a0<\/span><span style=\"font-weight: 400\">pdfrw<\/span><\/li>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">The Github page for\u00a0<\/span><span style=\"font-weight: 400\">PDFMiner<\/span><\/li>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">Using PyPDF2 for <\/span><span style=\"font-weight: 400\">Working with PDF files in Python<\/span><\/li>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">Working with PDF and Word Documents<\/span><\/li>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">The answer to StackOverflow question &#8211; <\/span><span style=\"font-weight: 400\">How to e<\/span><span style=\"font-weight: 400\">x<\/span><span style=\"font-weight: 400\">tract table as text from the PDF using Python?<\/span><\/li>\n<\/ul>\n<p dir=\"ltr\"><span style=\"font-weight: 400\">So overall, you need to understand that the PyPDF2 package is fast and pretty useful. It can be used for automating large jobs and using its capabilities for doing the job better.<\/span><\/p>\n<h3><span class=\"ez-toc-section\" id=\"final-thoughts\"><\/span>Final Thoughts<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p dir=\"ltr\">By taking a <a href=\"https:\/\/www.digitalvidya.com\/python-course\/\">Python Programming course<\/a> you can become a Python coding language master and a very skilled Python programmer. Any aspiring programmer can learn from Python&#8217;s basics and proceed after the course to finesse Python.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Python is a high-level language expressed with a simple syntax. This makes learning easy for new programmers. Some Python libraries can handle unstructured sources of data such as PDFs. Useful information such as audio, video, connections, buttons, business logic, and form fields can be found in PDFs. For displaying and sharing files, PDF or Portable [&hellip;]<\/p>\n","protected":false},"author":16370,"featured_media":165679,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":""},"categories":[11144],"tags":[11299,11300,11301,11302,11303,11298],"class_list":["post-164776","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-data-science","tag-pypdf2","tag-how-to-read-pdf-file-in-python","tag-read-pdf-in-python","tag-pypdf","tag-extract-text-from-pdf-python","tag-python-pdf"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.digitalvidya.com\/blog\/wp-json\/wp\/v2\/posts\/164776","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.digitalvidya.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.digitalvidya.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.digitalvidya.com\/blog\/wp-json\/wp\/v2\/users\/16370"}],"replies":[{"embeddable":true,"href":"https:\/\/www.digitalvidya.com\/blog\/wp-json\/wp\/v2\/comments?post=164776"}],"version-history":[{"count":0,"href":"https:\/\/www.digitalvidya.com\/blog\/wp-json\/wp\/v2\/posts\/164776\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.digitalvidya.com\/blog\/wp-json\/wp\/v2\/media\/165679"}],"wp:attachment":[{"href":"https:\/\/www.digitalvidya.com\/blog\/wp-json\/wp\/v2\/media?parent=164776"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.digitalvidya.com\/blog\/wp-json\/wp\/v2\/categories?post=164776"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.digitalvidya.com\/blog\/wp-json\/wp\/v2\/tags?post=164776"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}