|
xml:Proof <a schema for the rest of us/> |
v.02.06.10 Beta Thomas Sawyer (c)2002 |
A standard extensible and potable data language is extremely important to the IT community, thus the importance of XML technology, and its oft mention, as it has become that defacto standard in this regard. Yet it is widely held that XML is a bulky, less than optimal, implementation of such a standard. Fortunately there are ways in which the community itself can go about improving XML. xml:Proof is, in part, such an improvement.
XML, in and of itself, is simply a general data/metadata format --a way to organize data such that both the content and description of that content are bound together. But in itself it does not dictate the validity of that data. To patch this "hole" in XML, DTD or the Document Type Definition was made part of the XML specification. DTD has advantages. It is actually broader in applicability as it's syntax is not XML, but a superset, SGML. Yet this is also its disadvantage. The optimal solution would use XML itself as the base syntax, so that the same tools can be utilized for both the data/metadata markup and the validity markup. This is where schemas come into play. Schemas are XML document validity definitions, just as DTDs are, but they keep to the boundries of XML itself, i.e. schemas are marked-up with XML.
There are a number of schemas already available for XML, like TREX, RELAX, RELAX-NG, and Schematron. Offically the W3C has offered up their own XML-Schema. Should you place examples of all of these schemas side-by-side, along with an example of xml:Proof, xml:Proof will immediately distinguish itself from the rest. This is due to the fact that xml:Proof, unlike the others, actually utilizes the very tag names it intends to formalize, rather then invent a whole new set of its own. In fact xml:Proof has only two specially defined elements, the root tag and the arbit tag. As you can imagine this makes xml:Proof mark-up rather trivial to read and write. Additionly xml:Proof manages to do with so few speciality tags and attributes because it utilizes existing standard technologies to do much of its dirty work, that is Regular Expressions. Regular Expressions are well battle tested in the field, and there is little good reason to reinvent the wheel. Regular Expressions are a schema, using a broader sense of the word, in their own right, applicable to strings of text. As there are plenty of strings of text in XML documents, it isn't too hard to see how this might be useful. xml:Proof intends useage of Regular Expression insofar as is applicable in the context of XML. Utilizing this well known pre-existing technology, among its other features, XMProof is able to offer a unique and powerful schema to the XML community.
Personally I hate file extensions. Why file systems do not include a place for this description as they do for
the file name and last modified date is beyond me. I tend to blame MS-DOS. Oh well.
The extension for xml:Proof proofsheets, as they are called, is .xps.
xml:Proof is fully namespace aware, both in functionality and in application to an XML Document.
This requires further explination. Namespace prefixes serve as mere proxies to actual namespaces.
So while any arbitrary prefix can be used, a namespace itself, i.e. the uri, must be unique and persistent.
The namespace uri for xml:Proof is http://www.transami.net/namespace/xmlproof.
This namespace must be used on all of xml:Proof's special tags in order for any xml:Proof processor to function.
Further, when creating xml:Proof proofsheets, the namespaces of the elements and attributes being described must also be
taken into consideration with regards to the target XML document's. The elements and attributes of the XML document,
in other words, must partake of the same namespaces as their counterparts within the proofsheet.
This will become clearer as you read the rest of this document.
There are only two special tags in xml:Proof.
The first is the <proofsheet> tag. It is the root element of any xml:Proof schema document,
i.e. the proofsheet. The special root tag can take the alternate form of <schema>.
Both serve the same purpose.
The second special tag is the <arbit> tag. This tag is used to indicate an arbitrary
location in the XML document. It has a single valid attribute, xpath, which specifies the
the matching XML document nodes to which its die corresponds (see below).
Both of these special element tags and the special attrribute should always be prefixed with
reference to the xml:Proof namespace. While any arbitrary, but valid, prefix will do,
it is recommended that you use xp: for consistancy and clearity.
Witch exception to the special tags, all other tag and attribute names of an xml:Proof proofsheet are the same as those of the target XML documents it intends to model. The hiearchy of those elements are also the same. Thus the proofsheet is nearly as readable as any applicable target document. The text, or content, of elements and attributes is, in xml:Proof nomeclature, called a die. It may also be refered to as a cast and the act of writing or applying them, casting. A die consists of the following optional markers seperated by spaces:
=name=
The name is an identitier which names the die
such that it can be reused later in the proofsheet. This
provides a convenient means of die reuse. An element or attribute having
only this marker and no other will gain its die characteristics from any other die
identically named which has other markers within its die.
/regular expression/
The regular expression marker is dictates that the content
of an element or attribute must match against it to be considered valid.
The regular expression of a die effectively defaults to .* if excluded.
:datatype:
The datatype name is actually arbitrary, and can be anything desired.
xml:Proof itself dosen't care, but the utilization of an xml:Proof processor will!
Any given xml:Proof processor will generally "understand" the majority of common datatypes
and thus is able to typecast XML content into its underlying language of implementation.
Such is the main intent of datatype names in addition to validating content in simular fashion
to regular expressions. A good xml:Proof processor should also provide a means to add and alter
its internally recogzined datatypes. Any datatype it does not recognize will be treated as
CDATA, otherwise known as string or text.
@order@
The value of order may be tag, content-a..z, content-z..a>, or none.
If tag then all child elements of the casted element must appear in sequence as given within the proofsheet.
if content-a..z or content-z..a, then the content of all child elements of the casted element
must appear in alpahnumrical order, descending or ascending, respectively. The value none specifies that
no specific sort order is required and is the default if the marker is not given within the die.
Keep in mind this marker does not specify that each of the children elements must occur or that
one and only one of said children may appear. Rather, it only specifies that, should they appear,
they do so in the given order. An element thus cast must have child elements.
This marker is not applicable to attributes and will be ignored if used thus.
+closure+
The value of closure can be inclusive, exclusive, or none.
Inclusivity means that all child elements of the cast element must appear as given in the proofsheet,
but other elements may appear as thier siblings. Exclusivity means that all child elements of the cast element
must appear as given in the proofsheet and that no other elements may appear as thier siblings.
If this marker is not present within the die, the default value of none is assumed, which
relinquishes any neccessary closure on an elements child elemets. An element thus cast as
inclusive or exclusive must have children elements.
This marker is not applicable to attributes and will be ignored if used thus.
#range#
Specifies a range of how many of a given element or attribute may appear.
For elements, a valid range can be m..n or m...n,
inclusive and exclusive of n, respectively,
where m and n are unsigned integers
and m < n. This notation was borrowed from the Ruby programming language.
There is also the special case m..* (same as m...*) which of course means unbound.
0..* is the default, meaning none or any number of the element may appear within the document.
For attributes, only 0..1 and 1..1 are valid, as an attribute may appear no more
than once in any given element, with 0..1 being the default.
?option1,option2,...?
Where optionN is set to an arbitrary group name.
This option name defines an option group to which the the element belongs.
This specifies that one and only one of the elements sharing the same option group name
may appear within the target document. This can provide interesting relationships in that
elements and/or attributes having the same group names do not need to belong to the same parent!
But be warned: this can create a contridiction should an ancestor and one of its
children partake of the same group. Do not do this as it will render your documents
invalid by definition.
!collection!
Where collection is set to an arbitrary collection name.
This collection name defines a collective group to which the element belongs,
and specifies that all of suich elements and/or attrributes must appear together
within the document. Any given attribute or element can only belong to a single collection.
*track*
This is a special marker which does not dictate structure or content. It has a special purpose for XML datastores, like that implemented in DBXML. It specifies that this element should be specifically indexed. Tracking of particular XML elements in a datastore allows for fast search and retirieval, and more importantly fast aggregate functions to be applied to their values.
Here's an example of a die:
<Nword> =nword= #1..2# /^N/ :varchar: </Nword>
This die defines an xml tag named "Nword" to be any varchar beginning with the letter N and ooccuring only once or twice.
We have mentioned above xml:Proof's use of namespaces. In fact they are so fundemental, xml:Proof offers a variant notation for namespace declarations differing from the one recommened by the W3C. The W3C's recommendation is rather nebulous and clumsy, and further, clutters and obscures the information of relevance in an XML document. Therefore namespace declarations can be defined by document level processing instructions instead of within general element tags. Because XML processing instructions can be freely defined we have not violated any of the W3C standard by doing this, yet we have made our lives much improved!* This notation actually peacefully coexists with the standrard notation because it in effect does nothing more then insert a document level ATTRLIST for the namespaces defined.
Here is the top of an XML document using this alternate notation:
<?xml version="1.0" encoding="ISO-8859-1"?> <?xml:ns prefix="example" uri="http://www.transami.net/namespace/testing"?>This is effectively translated by the XML Processor into:
<?xml version="1.0" encoding="ISO-8859-1"?>
<!DocType docname [
<!ATTLIST docname xmlns:example 'http://www.transami.net/namespace/testing' CDATA>
]>
Thus the processing instruction xml:ns defines a namespace. The prefix and uri
attributes can also be labeled name and space, respectively. Subsequently any tag
or attribute prefixed with the prefix or name value will
thus be associated to this declared namespace.
Obviously, to date, XML Processors generally do not support this processing instruction, but it is hoped that this alternate notation will catch on in the XML community and will be generally adopted as a new standard. In the mean time all xml:Proof processors should provide a means to convert between the two different notations.
Schema declarations are similar to namespace declarations. They are declared via processing instructions as well. For example:
uri attribute, or its synonym space, define the kind of schema that is being utilized.
This is the specific namespace uri as defined by the schema's designers. In the case of xml:Proof, it
is "http://www.transami.net/namespace/xmlproof". It would be another string for, say, RELAX-NG or Schematron.
The url attribute, or its synonym source is a path name to
the .xps file. In this example case, it is a local file in the same location as the XML document itself.
This is neccessary since proofsheets can not be embedded in the document like DTDs can.
Interestingly, more than one schema can be declared. In so doing, schema declarations appearing higher
in the document have precedence over the later. This allows for a means of cast overiding. In our example,
for any given tag within the document, a matching die will first be searched for in example1.xps.
Only if it is not found there will example2.xps be searched. This can be quite useful in using borrowed
schemas. You can add new entries or overide existing entries without actually changing the original's.
*Note: In fact one rule has been violated: the reserved use of an instruction name matching /^xml/i. well, :-p
First let us look at a "traditional", "simple" XML-Schema example:
Now here's the near equivalent in xml:Proof, with a little extra added to show-off:
Notice the difference in the way namespaces are used. XML-Schema has its own namespace for every tag, seperate from the XML document's, which makes sense, since it uses its own set of tag names. Furthermore the document itself is forced to use an "instance" of the schema as the namespace of its elements and attributes. Thus the document is "confined" to the schema. xml:Proof on the other hand uses the same tag and attribute names as the document itself and thus the same freely defined namespace. Using the same namespace gives the two sets of data a greater association, without the limitations imposed by XML-Schema, and, last but certainly not least, is far easier to comprehend.
So all this is well and fine, but how does xml:Proof actually work? Well, that is farily simple really.
xml:Proof simply matches XPaths between the proofsheet and the document sharing the same namespace,
such that a particular die is applied to any corresponding document element or attribute. From the example given above,
you'll notice that the item element appears twice within the XML document. These two elements
match to the single proofsheet element of the same name. For instance the absolute XPath,
example:shiporder/item/quantity, containing 1 in the document, matches the same
absolute XPath, example:shiporder/item/quantity, containing =use_again= :unsigned: in the proofsheet.
This points out an important restriction to proofsheets: any possible absolute XPath within a proofsheet should only
be accounted for once.*
Arbitrary dies, cast via the <arbit> tag, overlap in applicability with the general absolute dies.
Thus if an element or attribute in a target XML document matches against an absolute XPath in the proofsheet and also
matches against an arbitrary XPath, it must conform to both dies. Further arbitrary dies themselves may overlap in
applicability.
*Note: If this is not adhered to it is not likely to cause a problem. The first occurance of a die will be matched and that will be that.
The XMLProof/Ruby API is a Ruby library for using xml:Proof. You can find documentation for its use here: xml:Proof/Ruby API Documentation.
xml:Proof, like all other schemas, is not a cure all for schema definition. It has its strengths and weaknesses. But no other schema, of which we are aware, matches its capabilites or ease of use. In the end, we believe, and we hope others will agree, xml:Proof is by far and away a better way to schema XML. It solves the majority of the requirements of a schematic meta-language while minimizing the complexity assocciated with them. Best of all it won't give you headaches.
<p:a>
<b/>
</p:a>
<b> inherits the prefix p from <p:a>.
Further, the root tag of a document, without a given prefix, would inherit the prefix of the first appearing namespace. Thus, with this new notation, there is no such beast called the default namespace. All tags and attributes, in the same fashion, either have a prefix or inherit one. The only exception is when no namespaces are declared. In this case all tags and attributes, "erroneously" prefixed or not belong to the null-namespace, or empty-namespace. Effectively this means no namespace. The null-namespace can be referenced by a prefix by setting the namespace uri to an empty string.
Dosen't this just make more sense? This seems so appealing to me that I almost made this a requirement of xml:Proof!. Oh well, the W3C keeps us working hard.