Класс DOMDocument

(PHP 5)

Введение

Представляет все содержимое HTML или XML документа; служит в качестве корня дерева документа.

Обзор классов

DOMDocument extends DOMNode {

/* Свойства */

readonly public string $actualEncoding ;

readonly public DOMConfiguration $config ;

readonly public DOMDocumentType $doctype ;

readonly public DOMElement $documentElement ;

public string $documentURI ;

public string $encoding ;

public bool $formatOutput ;

readonly public DOMImplementation $implementation ;

public bool $preserveWhiteSpace = true ;

public bool $recover ;

public bool $resolveExternals ;

public bool $standalone ;

public bool $strictErrorChecking = true ;

public bool $substituteEntities ;

public bool $validateOnParse = false ;

public string $version ;

readonly public string $xmlEncoding ;

public bool $xmlStandalone ;

public string $xmlVersion ;

/* Методы */

public __construct ([ string $version [, string $encoding ]] )

public DOMAttr createAttribute ( string $name )

public DOMAttr createAttributeNS ( string $namespaceURI , string $qualifiedName )

public DOMCDATASection createCDATASection ( string $data )

public DOMComment createComment ( string $data )

public DOMDocumentFragment createDocumentFragment ( void )

public DOMElement createElement ( string $name [, string $value ] )

public DOMElement createElementNS ( string $namespaceURI , string $qualifiedName [, string $value ] )

public DOMEntityReference createEntityReference ( string $name )

public DOMProcessingInstruction createProcessingInstruction ( string $target [, string $data ] )

public DOMText createTextNode ( string $content )

public DOMElement getElementById ( string $elementId )

public DOMNodeList getElementsByTagName ( string $name )

public DOMNodeList getElementsByTagNameNS ( string $namespaceURI , string $localName )

public DOMNode importNode ( DOMNode $importedNode [, bool $deep ] )

public mixed load ( string $filename [, int $options = 0 ] )

public bool loadHTML ( string $source [, int $options = 0 ] )

public bool loadHTMLFile ( string $filename [, int $options = 0 ] )

public mixed loadXML ( string $source [, int $options = 0 ] )

public void normalizeDocument ( void )

public bool registerNodeClass ( string $baseclass , string $extendedclass )

public bool relaxNGValidate ( string $filename )

public bool relaxNGValidateSource ( string $source )

public int save ( string $filename [, int $options ] )

public string saveHTML ([ DOMNode $node = NULL ] )

public int saveHTMLFile ( string $filename )

public string saveXML ([ DOMNode $node [, int $options ]] )

public bool schemaValidate ( string $filename )

public bool schemaValidateSource ( string $source )

public bool validate ( void )

public int xinclude ([ int $options ] )

/* Наследуемые методы */

public DOMNode DOMNode::appendChild ( DOMNode $newnode )

public string DOMNode::C14N ([ bool $exclusive [, bool $with_comments [, array $xpath [, array $ns_prefixes ]]]] )

public int DOMNode::C14NFile ( string $uri [, bool $exclusive [, bool $with_comments [, array $xpath [, array $ns_prefixes ]]]] )

public DOMNode DOMNode::cloneNode ([ bool $deep ] )

public int DOMNode::getLineNo ( void )

public string DOMNode::getNodePath ( void )

public bool DOMNode::hasAttributes ( void )

public bool DOMNode::hasChildNodes ( void )

public DOMNode DOMNode::insertBefore ( DOMNode $newnode [, DOMNode $refnode ] )

public bool DOMNode::isDefaultNamespace ( string $namespaceURI )

public bool DOMNode::isSameNode ( DOMNode $node )

public bool DOMNode::isSupported ( string $feature , string $version )

public string DOMNode::lookupNamespaceURI ( string $prefix )

public string DOMNode::lookupPrefix ( string $namespaceURI )

public void DOMNode::normalize ( void )

public DOMNode DOMNode::removeChild ( DOMNode $oldnode )

public DOMNode DOMNode::replaceChild ( DOMNode $newnode , DOMNode $oldnode )

}

Свойства

actualEncoding: Устарело. Кодировка документа, доступный только для чтения аналог encoding.
config: Устарело. Конфигурация использованная при вызове DOMDocument::normalizeDocument().
doctype: Объявление типа документа, соответствующее этому документу.
documentElement: Удобный атрибут, предоставляющий прямой доступ к узлу-потомку, как к элементу документа.
documentURI: Расположение документа или NULL, если не определено.
encoding: Кодировка документа, как она задана в объявлении XML. Этот атрибут отсутствует в итоговой DOM Level 3 спецификации, но это единственный путь для управления кодировкой XML документа в данной реализации.
formatOutput: Форматирует вывод, добавляя отступы и дополнительные пробелы.
implementation: Объект класса DOMImplementation, обрабатывающий этот документ.
preserveWhiteSpace: Указание не убирать лишние пробелы и отступы. По умолчанию TRUE.
recover: Патентованное свойство. Включает режим восстановления, то есть пытается разобрать некорректно составленные документы. Этот атрибут не входит в спецификацию DOM и является особенностью libxml.
resolveExternals: Установите в TRUE для загрузки внешних элементов из объявления типа документа. Может быть полезным при включении элементов с символьными данными в XML документ.
standalone: Устарело. Указание, что документ не зависит от других XML документов. Это можно определить из XML объявления. Свойство связано с xmlStandalone.
strictErrorChecking: Выбрасывает исключение DOMException при ошибке. По умолчанию TRUE.
substituteEntities: Патентованное свойство. Указывает, заменять или нет элементы документа. Этот атрибут не входит в спецификацию DOM и является особенностью libxml.
validateOnParse: Загружает DTD и проверяет документ на соответствие. По умолчанию FALSE.
version: Устарело. Версия XML, связанная с xmlVersion.
xmlEncoding: Атрибут задает, равно как и XML объявление, кодировку документа. Имеет значение NULL в случаях, когда атрибут не задан, либо значение неизвестно, если, например, документ создан в памяти.
xmlStandalone: Атрибут указывает, равно как и XML объявление, на то, что документ не зависит от других документов. Принимает значение FALSE, если не задан.
xmlVersion: Атрибут задает, равно как и XML объявление, версию документа. Если XML объявления в документе нет, но есть поддержка всех особенностей "XML", значение атрибута принимается равным "1.0".

Примечания

Замечание:
Расширение DOM использует кодировку UTF-8. Используйте функции utf8_encode() и utf8_decode() для работы с текстами в кодировке ISO-8859-1, либо Iconv в других кодировках.

Смотрите также

» W3C спецификация для Document

Содержание

DOMDocument::__construct — Создание нового DOMDocument объекта
DOMDocument::createAttribute — Создает новый атрибут
DOMDocument::createAttributeNS — Создает новый узел-атрибут с соответствующим ему пространством имен
DOMDocument::createCDATASection — Создает новый cdata узел
DOMDocument::createComment — Создает новый узел-комментарий
DOMDocument::createDocumentFragment — Создание фрагмента докуента
DOMDocument::createElement — Создает новый узел-элемент
DOMDocument::createElementNS — Создание нового узла-элемента с соответствующим пространством имен
DOMDocument::createEntityReference — Создание нового узла-ссылки на сущность
DOMDocument::createProcessingInstruction — Создает новый PI-узел
DOMDocument::createTextNode — Создает новый текстовый узел
DOMDocument::getElementById — Ищет элемент с заданным id
DOMDocument::getElementsByTagName — Ищет все элементы с заданным локальным именем
DOMDocument::getElementsByTagNameNS — Ищет элементы с заданным именем в определенном пространстве имен
DOMDocument::importNode — Импорт узла в текущий документ
DOMDocument::load — Загрузка XML из файла
DOMDocument::loadHTML — Загрузка HTML из строки
DOMDocument::loadHTMLFile — Загрузка HTML из файла
DOMDocument::loadXML — Загрузка XML из строки
DOMDocument::normalizeDocument — Нормализует документ
DOMDocument::registerNodeClass — Регистрация расширенного класса, используемого для создания базового типа узлов
DOMDocument::relaxNGValidate — Производит проверку документа на правильность построения посредством relaxNG
DOMDocument::relaxNGValidateSource — Проверяет документ посредством relaxNG
DOMDocument::save — Сохраняет XML дерево из внутреннего представления в файл
DOMDocument::saveHTML — Сохраняет документ из внутреннего представления в строку, используя HTML форматирование
DOMDocument::saveHTMLFile — Сохраняет документ из внутреннего представления в файл, используя HTML форматирование
DOMDocument::saveXML — Сохраняет XML дерево из внутреннего представления в виде строки
DOMDocument::schemaValidate — Проверяет действительности документа, основываясь на заданной схеме
DOMDocument::schemaValidateSource — Проверяет действительность документа, основываясь на схеме
DOMDocument::validate — Проверяет документ на соответствие его DTD
DOMDocument::xinclude — Проводит вставку XInclude разделов в объектах DOMDocument

Коментарии

Apr 11

Автор: Fernando H


Showing a quick example of how to use this class, just so that new users can get a quick start without having to figure it all out by themself. ( At the day of posting, this documentation just got added and is lacking examples. )



<?php



// Set the content type to be XML, so that the browser will   recognise it as XML.

header( "content-type: application/xml; charset=ISO-8859-15" );



// "Create" the document.

$xml = new DOMDocument( "1.0", "ISO-8859-15" );



// Create some elements.

$xml_album = $xml->createElement( "Album" );

$xml_track = $xml->createElement( "Track", "The ninth symphony" );



// Set the attributes.

$xml_track->setAttribute( "length", "0:01:15" );

$xml_track->setAttribute( "bitrate", "64kb/s" );

$xml_track->setAttribute( "channels", "2" );



// Create another element, just to show you can add any (realistic to computer) number of sublevels.

$xml_note = $xml->createElement( "Note", "The last symphony composed by Ludwig van Beethoven." );



// Append the whole bunch.

$xml_track->appendChild( $xml_note );

$xml_album->appendChild( $xml_track );



// Repeat the above with some different values..

$xml_track = $xml->createElement( "Track", "Highway Blues" );



$xml_track->setAttribute( "length", "0:01:33" );

$xml_track->setAttribute( "bitrate", "64kb/s" );

$xml_track->setAttribute( "channels", "2" );

$xml_album->appendChild( $xml_track );



$xml->appendChild( $xml_album );



// Parse the XML.

print $xml->saveXML();



?>



Output:

<Album>

  <Track length="0:01:15" bitrate="64kb/s" channels="2">

    The ninth symphony

    <Note>

      The last symphony composed by Ludwig van Beethoven.

    </Note>

  </Track>

  <Track length="0:01:33" bitrate="64kb/s" channels="2">Highway Blues</Track>

</Album>



If you want your PHP->DOM code to run under the .xml extension, you should set your webserver up to run the .xml extension with PHP ( Refer to the installation/configuration configuration for PHP on how to do this ).



Note that this:

<?php

$xml = new DOMDocument( "1.0", "ISO-8859-15" );

$xml_album = $xml->createElement( "Album" );

$xml_track = $xml->createElement( "Track" );

$xml_album->appendChild( $xml_track );

$xml->appendChild( $xml_album );

?>



is NOT the same as this:

<?php

// Will NOT work.

$xml = new DOMDocument( "1.0", "ISO-8859-15" );

$xml_album = new DOMElement( "Album" );

$xml_track = new DOMElement( "Track" );

$xml_album->appendChild( $xml_track );

$xml->appendChild( $xml_album );

?>



although this will work:

<?php

$xml = new DOMDocument( "1.0", "ISO-8859-15" );

$xml_album = new DOMElement( "Album" );

$xml->appendChild( $xml_album );

?>

2008-04-11 03:48:01

http://php5.kiev.ua/manual/ru/class.domdocument.html

May 23

Автор: cmyk777 at gmail dot com


This function may help to debug current dom element:



<?php

function dom_dump($obj) {

    if ($classname = get_class($obj)) {

        $retval = "Instance of $classname, node list: \n";

        switch (true) {

            case ($obj instanceof DOMDocument):

                $retval .= "XPath: {$obj->getNodePath()}\n".$obj->saveXML($obj);

                break;

            case ($obj instanceof DOMElement):

                $retval .= "XPath: {$obj->getNodePath()}\n".$obj->ownerDocument->saveXML($obj);

                break;

            case ($obj instanceof DOMAttr):

                $retval .= "XPath: {$obj->getNodePath()}\n".$obj->ownerDocument->saveXML($obj);

                //$retval .= $obj->ownerDocument->saveXML($obj);

                break;

            case ($obj instanceof DOMNodeList):

                for ($i = 0; $i < $obj->length; $i++) {

                    $retval .= "Item #$i, XPath: {$obj->item($i)->getNodePath()}\n".

"{$obj->item($i)->ownerDocument->saveXML($obj->item($i))}\n";

                }

                break;

            default:

                return "Instance of unknown class";

        }

    } else {

        return 'no elements...';

    }

    return htmlspecialchars($retval);

}

?>



Example usage:



<?php

$dom = new DomDocument();

$dom->load('test.xml');

$body = $dom->documentElement->getElementsByTagName('book');

echo '<pre>'.dom_dump($body).'<pre>';

?>



Output:



Instance of DOMNodeList, node list: 

Item #0, XPath: /library/book[1]

<book isbn="0345342968">

<title>Fahrenheit 451</title>

<author>R. Bradbury</author>

<publisher>Del Rey</publisher>

</book>

Item #1, XPath: /library/book[2]

<book isbn="0048231398">

<title>The Silmarillion</title>

<author>J.R.R. Tolkien</author>

<publisher>G. Allen &amp; Unwin</publisher>

</book>

Item #2, XPath: /library/book[3]

<book isbn="0451524934">

<title>1984</title>

<author>G. Orwell</author>

<publisher>Signet</publisher>

</book>

Item #3, XPath: /library/book[4]

<book isbn="031219126X">

<title>Frankenstein</title>

<author>M. Shelley</author>

<publisher>Bedford</publisher>

</book>

Item #4, XPath: /library/book[5]

<book isbn="0312863551">

<title>The Moon Is a Harsh Mistress</title>

<author>R. A. Heinlein</author>

<publisher>Orb</publisher>

</book>

2009-05-23 15:31:31

http://php5.kiev.ua/manual/ru/class.domdocument.html

Oct 31

Автор: fcartegnie


Be careful with formatOutput().



Creating an empty node like this:

createElement('foo','')

instead of

createElement('foo')

will break formatOutput.

2009-10-31 15:30:18

http://php5.kiev.ua/manual/ru/class.domdocument.html

Jan 27

Автор: jay at jaygilford dot com


Here's a small function I wrote to get all page links using the DOMDocument which will hopefully be of use to others



<?php

/**

 * @author Jay Gilford

 */

 

/**

 * get_links()

 * 

 * @param string $url

 * @return array

 */

function get_links($url) {

 

    // Create a new DOM Document to hold our webpage structure

    $xml = new DOMDocument();

 

    // Load the url's contents into the DOM

    $xml->loadHTMLFile($url);

 

    // Empty array to hold all links to return

    $links = array();

 

    //Loop through each <a> tag in the dom and add it to the link array

    foreach($xml->getElementsByTagName('a') as $link) {

        $links[] = array('url' => $link->getAttribute('href'), 'text' => $link->nodeValue);

    }

 

    //Return the links

    return $links;

}

?>

2010-01-27 10:46:16

http://php5.kiev.ua/manual/ru/class.domdocument.html

Feb 05

Автор: tloach at gmail dot com


For anyone else who has been having issues with formatOuput not working, here is a work-around:



rather than just doing something like:



<?php

$outXML = $xml->saveXML();

?>



force it to reload the XML from scratch, then it will format correctly:



<?php

$outXML = $xml->saveXML();

$xml = new DOMDocument();

$xml->preserveWhiteSpace = false;

$xml->formatOutput = true;

$xml->loadXML($outXML);

$outXML = $xml->saveXML();

?>

2010-02-05 12:01:16

http://php5.kiev.ua/manual/ru/class.domdocument.html

Mar 12

Автор: admin at beerpla dot net


After seeing many complaints about certain DOMDocument shortcomings, such as bad handling of encodings and always saving HTML fragments with <html>, <head>, and DOCTYPE, I decided that a better solution is needed.



So here it is: SmartDOMDocument. You can find it at http://beerpla.net/projects/smartdomdocument/



Currently, the main highlights are:



- SmartDOMDocument inherits from DOMDocument, so it's very easy to use - just declare an object of type SmartDOMDocument instead of DOMDocument and enjoy the new behavior on top of all existing functionality (see example below).



- saveHTMLExact() - DOMDocument has an extremely badly designed "feature" where if the HTML code you are loading does not contain <html> and <body> tags, it adds them automatically (yup, there are no flags to turn this behavior off).

Thus, when you call $doc->saveHTML(), your newly saved content now has <html><body> and DOCTYPE in it. Not very handy when trying to work with code fragments (XML has a similar problem).

SmartDOMDocument contains a new function called saveHTMLExact() which does exactly what you would want - it saves HTML without adding that extra garbage that DOMDocument does.



- encoding fix - DOMDocument notoriously doesn't handle encoding (at least UTF-8) correctly and garbles the output.

SmartDOMDocument tries to work around this problem by enhancing loadHTML() to deal with encoding correctly. This behavior is transparent to you - just use loadHTML() as you would normally.



- SmartDOMDocument Object As String - you can use a SmartDOMDocument object as a string which will print out its contents.

For example:

<?php

echo "Here is the HTML: $smart_dom_doc";

?>



I'm going to maintain this code and try to fix bugs as they come in.



Enjoy.

2010-03-12 04:12:02

http://php5.kiev.ua/manual/ru/class.domdocument.html

Nov 20

Автор: evert at er dot nl


A nice and simple node 2 array I wrote, worth a try ;) 



<?php

function getArray($node)

{

    $array = false;



    if ($node->hasAttributes())

    {

        foreach ($node->attributes as $attr)

        {

            $array[$attr->nodeName] = $attr->nodeValue;

        }

    }



    if ($node->hasChildNodes())

    {

        if ($node->childNodes->length == 1)

        {

            $array[$node->firstChild->nodeName] = $node->firstChild->nodeValue;

        }

        else

        {

            foreach ($node->childNodes as $childNode)

            {

                if ($childNode->nodeType != XML_TEXT_NODE)

                {

                    $array[$childNode->nodeName][] = $this->getArray($childNode);

                }

            }

        }

    }



    return $array;

}

?>

2010-11-20 04:17:48

http://php5.kiev.ua/manual/ru/class.domdocument.html

Jun 01

Автор: Nick M


You may need to save all or part of a DOMDocument as an XHTML-friendly string, something compliant with both XML and HTML 4. Here's the DOMDocument class extended with a saveXHTML method:



<?php



/**

 * XHTML Document

 *

 * Represents an entire XHTML DOM document; serves as the root of the document tree.

 */

class XHTMLDocument extends DOMDocument {



  /**

   * These tags must always self-terminate. Anything else must never self-terminate.

   * 

   * @var array

   */

  public $selfTerminate = array(

      'area','base','basefont','br','col','frame','hr','img','input','link','meta','param'

  );

  

  /**

   * saveXHTML

   *

   * Dumps the internal XML tree back into an XHTML-friendly string.

   *

   * @param DOMNode $node

   *         Use this parameter to output only a specific node rather than the entire document.

   */

  public function saveXHTML(DOMNode $node=null) {

    

    if (!$node) $node = $this->firstChild;

    

    $doc = new DOMDocument('1.0');

    $clone = $doc->importNode($node->cloneNode(false), true);

    $term = in_array(strtolower($clone->nodeName), $this->selfTerminate);

    $inner='';

    

    if (!$term) {

      $clone->appendChild(new DOMText(''));

      if ($node->childNodes) foreach ($node->childNodes as $child) {

        $inner .= $this->saveXHTML($child);

      }

    }

    

    $doc->appendChild($clone);

    $out = $doc->saveXML($clone);

    

    return $term ? substr($out, 0, -2) . ' />' : str_replace('><', ">$inner<", $out);



  }



}



?>



This hasn't been benchmarked, but is probably significantly slower than saveXML or saveHTML and should be used sparingly.

2011-06-01 13:16:09

http://php5.kiev.ua/manual/ru/class.domdocument.html

Jan 22

Автор: sites.sitesbr.net


How to objetify a DomDocument with hierarchy like:

<root>

    <item>

          <prop1>info1</prop1>

          <prop2>info2</prop2>

          <prop3>info3</prop3>

     </item>

    <item>

          <prop1>info1</prop1>

          <prop2>info2</prop2>

          <prop3>info3</prop3>

     </item>

</root>



It's possible to use in object style to retrieve information, as:



<?php

     $theNodeValue = $aitem->prop1;

?>



Here is the code: one Class and 2 functions.



<?php

 class ArrayNode{

       public $nodeName, $nodeValue;

 }



 function getChildNodeElements( $domNode ){

     $nodes = array();

     for( $i=0; $i < $domNode->childNodes->length; $i++){

       $cn = $domNode->childNodes->item($i);

       if( $cn->nodeType == 1){

           $nodes[] = $cn;

           }

     }

    return $nodes;

 }



 function getArrayNodes( $domDoc ){

     $res = array();



       for( $i=0; $i < $domDoc->childNodes->length; $i++){

       $cn = $domDoc->childNodes->item($i);

       # The first is the root tag...

          if( $cn->nodeType == 1){

               # But we want it's childNodes.

                $sub_cn = getChildNodeElements( $cn);

                # Found the tagName:

                $baseItemTagName = $sub_cn[0]->nodeName;

                break;

            }

        }



       $dnl = $domDoc->getElementsByTagName( $baseItemTagName);



       for( $i=0; $i< $dnl->length; $i++){

          $arrayNode = new ArrayNode();



      # Summary

      $arrayNode->nodeName = $dnl->item($i)->nodeName;

      $arrayNode->nodeValue = $dnl->item($i)->nodeValue;



      # Child Nodes

      $cn = $dnl->item($i)->childNodes;

      for( $k=0; $k<$cn->length; $k++){

           if( $cn->item($k)->nodeName == "#text" && trim($cn->item($k)->nodeValue) == "") continue;

           $arrayNode->{$cn->item($k)->nodeName} = $cn->item($k)->nodeValue;

      }



      # Attributes

      $attr = $dnl->item($i)->attributes;

      for( $k=0; $k < $attr->length; $k++){

           if(! is_null($attr)){

            if( $attr->item($k)->nodeName == "#text" && trim($attr->item($k)->nodeValue) == "") continue;

            $arrayNode->{$attr->item($k)->nodeName} = $attr->item($k)->nodeValue;

           }

      }



      $res[] = $arrayNode;



       }



     return $res;

 }

?>



To use it:



<?php



  # First you load a XML in a DomDocument variable.



   $url = "/path/to/yourxmlfile.xml";

   $domSrc = file_get_contents($url);

   $dom = new DomDocument();

   $dom->loadXML( $domSrc );



  # Then, you get the ArrayNodes from the DomDocument.



    $ans = getArrayNodes( $dom );



 

    for( $i=0; $i < count( $ans ) ; $i++){



    $cn =  $ans[ $i];



    $info1 =  $cn->prop1;

    $info2 =  $cn->prop2;

    $info3 =  $cn->prop3;

      

         // ...

 

   }



?>

2013-01-22 02:37:05

http://php5.kiev.ua/manual/ru/class.domdocument.html

Oct 28

Автор: danny dot nunez15 at gmail dot com


A simple function to grab all links in a page. 



    function get_links($url) {



        // Create a new DOM Document to hold our webpage structure 

        $xml = new DOMDocument();



        // Load the url's contents into the DOM 



        $xml->loadHTMLFile($url);



        // Empty array to hold all links to return 

        $links = array();



        //Loop through each <a> tag in the dom and add it to the link array 

        foreach ($xml->getElementsByTagName('a') as $link) {

            $url = $link->getAttribute('href');

            if (!empty($url)) {

                $links[] = $link->getAttribute('href');

            }

        }



        //Return the links 

        return $links;

    }

2013-10-28 16:54:33

http://php5.kiev.ua/manual/ru/class.domdocument.html

Nov 11

Автор: qrworld.net


In this post http://softontherocks.blogspot.com/2014/11/descargar-el-contenido-de-una-url_11.html I found a simple way to get the content of a URL with DOMDocument, loadHTMLFile and saveHTML().



function getURLContent($url){

    $doc = new DOMDocument;

    $doc->preserveWhiteSpace = FALSE;

    @$doc->loadHTMLFile($url);

    return $doc->saveHTML();

}

2014-11-11 18:35:33

http://php5.kiev.ua/manual/ru/class.domdocument.html

May 14

Автор: ingjetel at gmail dot com


Easy function for basic output of XML file via DOM parsing



<?php

$dom = new DomDocument();

$dom->load("./file.xml") or die("error");

$start = $dom->documentElement;

fc($start);



function fc($node) {

  $child = $node->childNodes;

  foreach($child as $item) {

    if ($item->nodeType == XML_TEXT_NODE) {

      if (strlen(trim($item->nodeValue))) echo trim($item->nodeValue)."<br/>";

    }

    else if ($item->nodeType == XML_ELEMENT_NODE) fc($item);

  }

}

?>

2015-05-14 02:54:39

http://php5.kiev.ua/manual/ru/class.domdocument.html

Dec 17

Автор: developer at nabtron dot com


For those landing here and checking for encoding issue with utf-8 characteres, it's pretty easy to correct it, without adding any additional output tag to your html.



We'll be utilizing: mb_convert_encoding



Thanks to the user who shared: SmartDOMDocument in previous comments, I got the idea of solving it. However I truly wish that he shared the method instead of giving a link.



Anyway coming back to the solution, you can simply use:



<?php



            // checks if the content we're receiving isn't empty, to avoid the warning

            if ( empty( $content ) ) {

                return false;

            }



            // converts all special characters to utf-8

            $content = mb_convert_encoding($content, 'HTML-ENTITIES', 'UTF-8');



            // creating new document

            $doc = new DOMDocument('1.0', 'utf-8');



            //turning off some errors

            libxml_use_internal_errors(true);



            // it loads the content without adding enclosing html/body tags and also the doctype declaration

            $doc->LoadHTML($content, LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD);



            // do whatever you want to do with this code now



?>



I hope it solves the issue for someone! If you need my help or service to fix your code, you can reach me on nabtron.com or contact me at the email mentioned with this comment.

2015-12-17 09:34:02

http://php5.kiev.ua/manual/ru/class.domdocument.html

Aug 22

Автор: biker dot mike at gmx dot com


Look out for the following gotcha when loading XML from a string:



<?php

$doc = new \DOMDocument;

$doc->documentURI = $myXmlFilename;

$doc->loadXML($myXmlString);

?>



documentURI is now set to the value of $myXmlFilename, right?



Wrong!



It's set to the current working directory.  If you want to manually set documentURI to something other than the CWD, do so AFTER the call to loadXML().



E.g.:

<?php

$doc = new \DOMDocument;

$doc->loadXML($myXmlString);

$doc->documentURI = $myXmlFilename;

?>



documentURI really is now set to the value of $myXmlFilename.

2016-08-22 16:36:32

http://php5.kiev.ua/manual/ru/class.domdocument.html

Jul 08

Автор: ashjkshdu283 at gmail dot com


/* Function evolved from jay at jaygilford dot com post

  * This function will return an array of the values of the specified

  * attribute ($attr) for all the Dom Document object's elements 

  */



<?php



function getAttrData(string $attr, DomDocument $dom) { 

    // Empty array to hold all classes to return 

    $attrData = array(); 



    //Loop through each tag in the dom and add it's attribute data to the array 

    foreach($dom->getElementsByTagName('*') as $tag) {

        if(empty($tag->getAttribute($attr)) === false) {

            array_push($attrData, $tag->getAttribute($attr));

        }

    } 



    //Return the array of attribute data

    return array_unique($attrData); 

}



$html = '

<!DOCTYPE html>

<html>

<head>

<title>Page Title</title>

</head>

<body>

<a href="#someLink" id="someLink" class="link-class">Some Link</a>

<a href="#someOtherLink" id="someOtherLink" class="link-class">Some Other Link</a>

<h1 id="header1" class="header-class">My First Heading</h1>

<p id="para1" class="para-class">My first paragraph.</p>

</body>

</html>';

$dom = new DOMDocument();

$dom->loadHtml($html);

$dom->saveHTML();

var_dump(getAttrData('class', $dom));

2018-07-08 02:39:25

http://php5.kiev.ua/manual/ru/class.domdocument.html

Mar 04

Автор: pastormontesinos at gmail dot com


For using safely with script nodes when parsing, best option is extending DOMDocument, keeping script tags while DOMDocument process and rearrange them just after saveHTML function is called. Here is my custom class.



<?php 



class SafeDOMDocument extends \DOMDocument

{

    const REGEX_JS            = '#(\s*<!--(\[if[^\n]*>)?\s*(<script.*</script>)+\s*(<!\[endif\])?-->)|(\s*<script.*</script>)#isU';

    const SUBSTITUTION_FORMAT = '<!--<script class="script_%s"></script>-->';

    private $matchedScripts = [];



    public function loadHTML($source, $options = 0)

    {

        $this->formatOutput        = false;

        $this->preserveWhiteSpace  = true;

        $this->validateOnParse     = false;

        $this->strictErrorChecking = false;

        $this->recover             = false;

        $this->resolveExternals    = false;

        $this->substituteEntities  = false;

        $matches = [];

        $success = preg_match_all(self::REGEX_JS, $source, $matches);



        if ($success && !empty($matches)) {

            foreach ($matches[0] as $match) {

                $storedScript = rtrim(ltrim($match, "\n\r\t "), "\n\r\t ");

                $scriptId = md5($storedScript);

                $key = sprintf(self::SUBSTITUTION_FORMAT, $scriptId);

                $source = str_replace($match, $key, $source);

                $this->matchedScripts[$key] = $storedScript;

            }

        }



        return parent::loadHTML($source, $options);

    }



    public function saveHTML(DOMNode $node = null)

    {

        $output = parent::saveHTML($node);



        if (count($this->matchedScripts)) {

            foreach ($this->matchedScripts as $substitution => $originalSnippet) {

                $output = str_replace($substitution, $originalSnippet, $output);

            }

        }



        return $output;

    }

}

?>

2021-03-04 12:34:13

http://php5.kiev.ua/manual/ru/class.domdocument.html

Nov 17

Автор: andreas at userbrain dot com


After struggling with parsing and modifying partial HTML content for several hours, I came to this solution which does work for me and is relatively simple compared to what else I found online.



This solution fixes unwanted DOCTYPE and html, body tags as well as encoding issues.



<?php



// Assumption: content is utf-8 encoded

$content = "<h1>This is a heading</h1><p>This is a paragraph</p>";



// Load content to a div and specify encoding with a meta tag

$temp_dom = new DOMDocument();

$temp_dom->loadHTML("<meta http-equiv='Content-Type' content='charset=utf-8' /><div>$content</div>");



// As loadHTML() adds a DOCTYPE as well as <html> and <body> tag, let’s create another DOMDocument and import just the nodes we want

$dom = new DOMDocument();

$first_div = $temp_dom->getElementsByTagName('div')[0];

$first_div_node = $dom->importNode($first_div, true);

$dom->appendChild($first_div_node);



// Do whatever you want to do

$dom->getElementsByTagName('h1')[0]->setAttribute('class', 'happy');



// You could also just echo $dom->saveHtml() if you don’t mind the div and whitespace 

echo substr(trim($dom->saveHtml()), 5, -6);



// Outputs: <h1 class="happy">This is a heading</h1><p>This is a paragraph</p>

?>

2021-11-17 15:16:42

http://php5.kiev.ua/manual/ru/class.domdocument.html

Mar 17

Автор: 610010559 at qq dot com


when you add the new element to formatted XML data through appendChild() method, you would the new element you add is not be formatted(that is not indexed, not line break).  here is my solution (in short load the xml without preserve white space, ), example show as below:

<?php

$doc = new \DOMDocument();

$doc->formatOutput = true;

$doc->preserveWhiteSpace = false;//that is key, default value is true. 

$doc->loadXML($xmlStr);

$doc->appendChild($doc->createElement('php', '666'))

$formattedXMLStr = $doc->saveXML();//DOMDocument wold format the xml str for you

echo $formattedXMlStr;

?>

it take me some time to try it out. hope it save your time.

2022-03-17 04:39:36

http://php5.kiev.ua/manual/ru/class.domdocument.html

Sep 25

Автор: devour at php dot net


While DOMDocument can technically be used to parse HTML, it is not ideal for HTML documents and is better suited for processing well-formed XML. One of the primary issues with using DOMDocument for HTML is its strict handling of special characters, such as the ampersand (&).



DOMDocument requires that ampersands be escaped as &amp;, which is in line with XML standards but can be counterintuitive for handling real-world HTML, where raw & characters are commonly found, especially in URLs and text. This behavior stems from the underlying XML-based parser (libxml), which treats HTML with the same strictness as XML.



This problem has been reported as far back as 2001, yet the same parsing errors continue to occur when using DOMDocument on HTML documents today.



A common workaround developers use is to suppress the error reporting from DOMDocument, particularly when parsing errors like unescaped ampersands occur. However, suppressing these errors is not recommended, especially in production environments, as it can hide important issues and pose potential security risks. Ignoring or suppressing errors can leave warnings unnoticed, which may result in vulnerabilities if not properly addressed.



For these reasons, it's advisable to use DOMDocument primarily for XML documents, or to consider more appropriate libraries  when working with HTML to avoid these issues.



theCoder / MV

2024-09-25 07:46:46

http://php5.kiev.ua/manual/ru/class.domdocument.html

DOMComment::__construct

DOMDocument::__construct

DOM

PHP Manual

PHP5

Для web разработчика

Dec 22
Класс DOMDocument

Класс DOMDocument

Введение

Обзор классов

Свойства

Примечания

Смотрите также

Содержание

Коментарии

PHP5

Для web разработчика

Dec 22Класс DOMDocument

Класс DOMDocument

Введение

Обзор классов

Свойства

Примечания

Смотрите также

Содержание

Коментарии

Dec 22
Класс DOMDocument