I developed this exact same tech earlier this year. It was a pain trying to analyze all the MANY different ways people can screw up html, but the success rate is pretty high across the net. I've been thinking of making it opensource since I no longer use it.