{"id":2316,"date":"2010-07-31T20:21:56","date_gmt":"2010-08-01T00:21:56","guid":{"rendered":"http:\/\/www.hoogervorst.ca\/arthur\/?p=2316"},"modified":"2010-07-31T20:23:01","modified_gmt":"2010-08-01T00:23:01","slug":"the-state-of-the-machine","status":"publish","type":"post","link":"http:\/\/www.hoogervorst.ca\/arthur\/?p=2316","title":{"rendered":"The State of the Machine"},"content":{"rendered":"<p><span class=\"dropcap\">I<\/span> get cranky when I see people use regular expressions or simple substring routines when extracting strings from, for example, e-mail addresses. The first method, while powerful, is memory hungry, the second method is plain <em>childish<\/em>. You should only use substring\/copy methods if you&#8217;re hundred percent certain that the data is formatted and well-formed (that is, it comes through exactly as you expect it to. In the case of e-mail addresses, this is of course, not true.  After all, e-mail addresses can come in any format. The following samples are all legal: &#8220;hey@you.com&#8221;, &#8220;hey@you.com (Hey You)&#8221;, &#8220;<hey@you.com> Hey You&#8221;, &#8220;Hey, You <hey@you.com>&#8220;. Your simple substring copy function would most likely have troubles resolving all of these e-mail variants.\n<\/p>\n<p>During my Roundabout tenure, we ran into issues where extraction of names\/e-mail from e-mail headers didn&#8217;t work out as originally planned. I was not surprised to find those evil substring routines in the code and literally rewrote that <a href=\"http:\/\/roundabout.cvs.sourceforge.net\/viewvc\/roundabout\/Phoenix%20Main\/ParserSup.pas?revision=1.5&#038;view=markup\">into a state machine<\/a> (look for <em>HeaderAddressToStringList<\/em>). Extremely elegant and very effective.\n<\/p>\n<p>Why use a state machine then? Because with string operations like this, looping through a string is a lot faster than trying hundreds of &#8220;if conditions&#8221; to cover all these e-mail cases. Keep in mind that simplicity is the key though: the more states you define, the complexer the code.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>I get cranky when I see people use regular expressions or simple substring routines when extracting strings from, for example, e-mail addresses. The first method, while powerful, is memory hungry, the second method is plain childish. You should only use &hellip; <a href=\"http:\/\/www.hoogervorst.ca\/arthur\/?p=2316\">Continue reading <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":[],"categories":[27],"tags":[146,86,176],"_links":{"self":[{"href":"http:\/\/www.hoogervorst.ca\/arthur\/index.php?rest_route=\/wp\/v2\/posts\/2316"}],"collection":[{"href":"http:\/\/www.hoogervorst.ca\/arthur\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"http:\/\/www.hoogervorst.ca\/arthur\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"http:\/\/www.hoogervorst.ca\/arthur\/index.php?rest_route=\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"http:\/\/www.hoogervorst.ca\/arthur\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2316"}],"version-history":[{"count":0,"href":"http:\/\/www.hoogervorst.ca\/arthur\/index.php?rest_route=\/wp\/v2\/posts\/2316\/revisions"}],"wp:attachment":[{"href":"http:\/\/www.hoogervorst.ca\/arthur\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2316"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"http:\/\/www.hoogervorst.ca\/arthur\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2316"},{"taxonomy":"post_tag","embeddable":true,"href":"http:\/\/www.hoogervorst.ca\/arthur\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2316"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}