xml:Proof
<a schema for the rest of us/>
v.02.06.10 Beta
 Thomas Sawyer (c)2002

Introduction

A standard extensible and potable data language is extremely important to the IT community, thus the importance of XML technology, and its oft mention, as it has become that defacto standard in this regard. Yet it is widely held that XML is a bulky, less than optimal, implementation of such a standard. Fortunately there are ways in which the community itself can go about improving XML. xml:Proof is, in part, such an improvement.

XML, in and of itself, is simply a general data/metadata format --a way to organize data such that both the content and description of that content are bound together. But in itself it does not dictate the validity of that data. To patch this "hole" in XML, DTD or the Document Type Definition was made part of the XML specification. DTD has advantages. It is actually broader in applicability as it's syntax is not XML, but a superset, SGML. Yet this is also its disadvantage. The optimal solution would use XML itself as the base syntax, so that the same tools can be utilized for both the data/metadata markup and the validity markup. This is where schemas come into play. Schemas are XML document validity definitions, just as DTDs are, but they keep to the boundries of XML itself, i.e. schemas are marked-up with XML.

There are a number of schemas already available for XML, like TREX, RELAX, RELAX-NG, and Schematron. Offically the W3C has offered up their own XML-Schema. Should you place examples of all of these schemas side-by-side, along with an example of xml:Proof, xml:Proof will immediately distinguish itself from the rest. This is due to the fact that xml:Proof, unlike the others, actually utilizes the very tag names it intends to formalize, rather then invent a whole new set of its own. In fact xml:Proof has only two specially defined elements, the root tag and the arbit tag. As you can imagine this makes xml:Proof mark-up rather trivial to read and write. Additionly xml:Proof manages to do with so few speciality tags and attributes because it utilizes existing standard technologies to do much of its dirty work, that is Regular Expressions. Regular Expressions are well battle tested in the field, and there is little good reason to reinvent the wheel. Regular Expressions are a schema, using a broader sense of the word, in their own right, applicable to strings of text. As there are plenty of strings of text in XML documents, it isn't too hard to see how this might be useful. xml:Proof intends useage of Regular Expression insofar as is applicable in the context of XML. Utilizing this well known pre-existing technology, among its other features, XMProof is able to offer a unique and powerful schema to the XML community.


Overview

File Extension and Namespace

Personally I hate file extensions. Why file systems do not include a place for this description as they do for the file name and last modified date is beyond me. I tend to blame MS-DOS. Oh well. The extension for xml:Proof proofsheets, as they are called, is .xps.

xml:Proof is fully namespace aware, both in functionality and in application to an XML Document. This requires further explination. Namespace prefixes serve as mere proxies to actual namespaces. So while any arbitrary prefix can be used, a namespace itself, i.e. the uri, must be unique and persistent. The namespace uri for xml:Proof is http://www.transami.net/namespace/xmlproof. This namespace must be used on all of xml:Proof's special tags in order for any xml:Proof processor to function. Further, when creating xml:Proof proofsheets, the namespaces of the elements and attributes being described must also be taken into consideration with regards to the target XML document's. The elements and attributes of the XML document, in other words, must partake of the same namespaces as their counterparts within the proofsheet. This will become clearer as you read the rest of this document.


Root and Arbit Tags

There are only two special tags in xml:Proof.

The first is the <proofsheet> tag. It is the root element of any xml:Proof schema document, i.e. the proofsheet. The special root tag can take the alternate form of <schema>. Both serve the same purpose.

The second special tag is the <arbit> tag. This tag is used to indicate an arbitrary location in the XML document. It has a single valid attribute, xpath, which specifies the the matching XML document nodes to which its die corresponds (see below).

Both of these special element tags and the special attrribute should always be prefixed with reference to the xml:Proof namespace. While any arbitrary, but valid, prefix will do, it is recommended that you use xp: for consistancy and clearity.


The Die is Cast

Witch exception to the special tags, all other tag and attribute names of an xml:Proof proofsheet are the same as those of the target XML documents it intends to model. The hiearchy of those elements are also the same. Thus the proofsheet is nearly as readable as any applicable target document. The text, or content, of elements and attributes is, in xml:Proof nomeclature, called a die. It may also be refered to as a cast and the act of writing or applying them, casting. A die consists of the following optional markers seperated by spaces:


Here's an example of a die:

	<Nword> =nword= #1..2# /^N/ :varchar: </Nword>

This die defines an xml tag named "Nword" to be any varchar beginning with the letter N and ooccuring only once or twice.


Namespaces and Schema Declerations

We have mentioned above xml:Proof's use of namespaces. In fact they are so fundemental, xml:Proof offers a variant notation for namespace declarations differing from the one recommened by the W3C. The W3C's recommendation is rather nebulous and clumsy, and further, clutters and obscures the information of relevance in an XML document. Therefore namespace declarations can be defined by document level processing instructions instead of within general element tags. Because XML processing instructions can be freely defined we have not violated any of the W3C standard by doing this, yet we have made our lives much improved!* This notation actually peacefully coexists with the standrard notation because it in effect does nothing more then insert a document level ATTRLIST for the namespaces defined.

Here is the top of an XML document using this alternate notation:

  <?xml version="1.0" encoding="ISO-8859-1"?>
  <?xml:ns prefix="example" uri="http://www.transami.net/namespace/testing"?>
This is effectively translated by the XML Processor into:
  <?xml version="1.0" encoding="ISO-8859-1"?>
  <!DocType docname [
    <!ATTLIST docname xmlns:example 'http://www.transami.net/namespace/testing' CDATA>
  ]>
Thus the processing instruction xml:ns defines a namespace. The prefix and uri attributes can also be labeled name and space, respectively. Subsequently any tag or attribute prefixed with the prefix or name value will thus be associated to this declared namespace.

Obviously, to date, XML Processors generally do not support this processing instruction, but it is hoped that this alternate notation will catch on in the XML community and will be generally adopted as a new standard. In the mean time all xml:Proof processors should provide a means to convert between the two different notations.

Schema declarations are similar to namespace declarations. They are declared via processing instructions as well. For example:

<?xml version="1.0" encoding="ISO-8859-1"?> <?xml:ns prefix="example" uri="http://www.transami.net/namespace/testing"?> <?xml:schema uri="http://www.transami.net/namespace/xmlproof" url="example1.xsp"?> <?xml:schema uri="http://www.transami.net/namespace/xmlproof" url="example2.xsp"?> The uri attribute, or its synonym space, define the kind of schema that is being utilized. This is the specific namespace uri as defined by the schema's designers. In the case of xml:Proof, it is "http://www.transami.net/namespace/xmlproof". It would be another string for, say, RELAX-NG or Schematron.

The url attribute, or its synonym source is a path name to the .xps file. In this example case, it is a local file in the same location as the XML document itself. This is neccessary since proofsheets can not be embedded in the document like DTDs can.

Interestingly, more than one schema can be declared. In so doing, schema declarations appearing higher in the document have precedence over the later. This allows for a means of cast overiding. In our example, for any given tag within the document, a matching die will first be searched for in example1.xps. Only if it is not found there will example2.xps be searched. This can be quite useful in using borrowed schemas. You can add new entries or overide existing entries without actually changing the original's.

*Note: In fact one rule has been violated: the reserved use of an instruction name matching /^xml/i. well, :-p


Example

First let us look at a "traditional", "simple" XML-Schema example:

<?xml version="1.0" encoding="ISO-8859-1"?> <shiporder orderid="889923" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="shiporder.xsd"> <orderperson>John Smith</orderperson> <shipto> <name>Ola Nordmann</name> <address>Langgt 23</address> <city>4000 Stavanger</city> <country>Norway</country> </shipto> <item> <title>Empire Burlesque</title> <note>Special Edition</note> <quantity>1</quantity> <price>10.90</price> </item> <item> <title>Hide your heart</title> <quantity>1</quantity> <price>9.90</price> </item> </shiporder> <?xml version="1.0" encoding="ISO-8859-1" ?> <xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema"> <xs:element name="shiporder"> <xs:complexType> <xs:sequence> <xs:element name="orderperson" type="xs:string"/> <xs:element name="shipto"> <xs:complexType> <xs:sequence> <xs:element name="name" type="xs:string"/> <xs:element name="address" type="xs:string"/> <xs:element name="city" type="xs:string"/> <xs:element name="country" type="xs:string"/> </xs:sequence> </xs:complexType> </xs:element> <xs:element name="item" maxOccurs="unbounded"> <xs:complexType> <xs:sequence> <xs:element name="title" type="xs:string"/> <xs:element name="note" type="xs:string" minOccurs="0"/> <xs:element name="quantity" type="xs:positiveInteger"/> <xs:element name="prize" type="xs:decimal"/> </xs:sequence> </xs:complexType> </xs:element> </xs:sequence> <xs:attribute name="orderid" type="xs:string" use="required"/> </xs:complexType> </xs:element> </xs:schema>

Now here's the near equivalent in xml:Proof, with a little extra added to show-off:

<?xml version="1.0" encoding="ISO-8859-1"?> <?xml:ns name="example" space="http://www.transami.net/namespace/testing"?> <?xml:schema source="example1.xsp" space="http://www.transami.net/namespace/xmlproof"?> <example:shiporder orderid="889923"> <orderperson>John Smith</orderperson> <shipto> <name>Ola Nordmann</name> <address>Langgt 23</address> <city>4000 Stavanger</city> <country>Norway</country> </shipto> <item> <title>Empire Burlesque</title> <note>Special Edition</note> <quantity>1</quantity> <price>10.90</price> </item> <item> <title>Hide your heart</title> <quantity>1</quantity> <price>9.90</price> </item> </example:shiporder> <?xml version="1.0" encoding="ISO-8859-1" ?> <?xml:ns name="example" space="http://www.transami.net/namespace/testing" ?> <?xml:ns name="xp" space="http://www.transami.net/namespace/xmlproof" ?> <xp:proofsheet> <example:shiporder orderid=":int:"> <orderperson> :text: ?bywho? </orderperson> <orderclerk> :text: ?bywho? </orderclerk> <shipto> #1..1# @true@ <name> :text: </name> <address> :text: </address> <city> :text: </city> <country> :text: </country> </shipto> <item> #1..*# @true@ <title> :text: </title> <note> :text: </note> <quantity> =use_again= :unsigned: </quantity> <overstock> =use_again= </overstock> <price> :float: </price> </item> </example:shiporder> </xp:proofsheet>

Notice the difference in the way namespaces are used. XML-Schema has its own namespace for every tag, seperate from the XML document's, which makes sense, since it uses its own set of tag names. Furthermore the document itself is forced to use an "instance" of the schema as the namespace of its elements and attributes. Thus the document is "confined" to the schema. xml:Proof on the other hand uses the same tag and attribute names as the document itself and thus the same freely defined namespace. Using the same namespace gives the two sets of data a greater association, without the limitations imposed by XML-Schema, and, last but certainly not least, is far easier to comprehend.


Functionality

So all this is well and fine, but how does xml:Proof actually work? Well, that is farily simple really. xml:Proof simply matches XPaths between the proofsheet and the document sharing the same namespace, such that a particular die is applied to any corresponding document element or attribute. From the example given above, you'll notice that the item element appears twice within the XML document. These two elements match to the single proofsheet element of the same name. For instance the absolute XPath, example:shiporder/item/quantity, containing 1 in the document, matches the same absolute XPath, example:shiporder/item/quantity, containing =use_again= :unsigned: in the proofsheet. This points out an important restriction to proofsheets: any possible absolute XPath within a proofsheet should only be accounted for once.*

Arbitrary dies, cast via the <arbit> tag, overlap in applicability with the general absolute dies. Thus if an element or attribute in a target XML document matches against an absolute XPath in the proofsheet and also matches against an arbitrary XPath, it must conform to both dies. Further arbitrary dies themselves may overlap in applicability.

*Note: If this is not adhered to it is not likely to cause a problem. The first occurance of a die will be matched and that will be that.


XMLProof/Ruby API

The XMLProof/Ruby API is a Ruby library for using xml:Proof. You can find documentation for its use here: xml:Proof/Ruby API Documentation.


Conclusion

xml:Proof, like all other schemas, is not a cure all for schema definition. It has its strengths and weaknesses. But no other schema, of which we are aware, matches its capabilites or ease of use. In the end, we believe, and we hope others will agree, xml:Proof is by far and away a better way to schema XML. It solves the majority of the requirements of a schematic meta-language while minimizing the complexity assocciated with them. Best of all it won't give you headaches.

  • xml:Proof, like all other schemas, is not a cure all for schema definition. It has its strengths and weaknesses. But we beleive that an analysis of the schematic problem set indicates that no other schema matches xml:Proofs capabilites or ease of use.
  • xml:Proof is a better way to schema XML. It solves the majority of the requirements of a schematic meta-language while minimizing the complexity assocciated with them.


  • After Thoughts

    Honestly I wish prefixes and namespaces were inherited, such that a non-prefixed tag inherits the prefix of its closest prefixed ancestor. Thus in the example:
      <p:a>
        <b/>
      </p:a>
    
    <b> inherits the prefix p from <p:a>.

    Further, the root tag of a document, without a given prefix, would inherit the prefix of the first appearing namespace. Thus, with this new notation, there is no such beast called the default namespace. All tags and attributes, in the same fashion, either have a prefix or inherit one. The only exception is when no namespaces are declared. In this case all tags and attributes, "erroneously" prefixed or not belong to the null-namespace, or empty-namespace. Effectively this means no namespace. The null-namespace can be referenced by a prefix by setting the namespace uri to an empty string.

    Dosen't this just make more sense? This seems so appealing to me that I almost made this a requirement of xml:Proof!. Oh well, the W3C keeps us working hard.