Tuesday, April 28, 2009
Monday, March 16, 2009
Since my last post I have been sent a few and read a few other articles that I thought might be worthy of note related to what clouds are and why it is important.
Sun's CTO Greg Papadopoulos did an interview with Informationweek and discussed some of the efforts that SUN has going on. One of the biggest items of note in that article is the new Drizzle product coming out of the MySQL work. There are many articles and blog posts on the topic as well but the gist is that it is an effort to Cloud-smarten MySQL. Seeing how "MySQL gets used, people only use a subset of the relational capabilities because it has a horizontal scale and there are all these concerns that go into large-scale deployments. So Drizzle sort of strips back down to a small core and then builds up the distributed capabilities." I have made no secret of my feeling that Data is no longer relational. So I think this is encouraging news though as seems to be the case lately with SUN it may be too late for MySQL.
(Aside... Greg Papadopoulos and David Douglas recently co-authored Citizen Engineer: A handbook for socially responsible engineering
As if they noticed Microsoft pushing cloud tech Amazon announced a few new items in their own cloud initiatives. The ability to reserve capacity for known needs. They also expanded the ability of Windows based services in another Amazon zone. All of this from a company that sells stuff, proof that necessity is the mother of invention.
Friday, May 16, 2008
Data is no longer relational
Ok, so the headline may be a little of an over statement for effect. Perhaps I could have said that data that people are interested in is no longer relational but then it wouldn't have been nearly so pithy.
With much respect to Edgar Codd and his invention of the relational model for database storage I think it is time to move forward. Relational databases are great for things like financial models, personnel data and other Enterprise Systems as well as many other standard, repetitive data. It gave a solid reference point to learn data structures and modeling to several generations of budding Computer Scientists.
When I say it is time to move forward I am not saying we should immediately move all data systems to Object Oriented Databases or otherwise induce data chaos. What I do want to push is the idea of unstructured data. Computers are great with rules, with structure and with fundamentally binary relationships. Algorithms are starting to mature around unstructured data (for an example go search google.) but it is still not widespread or well understood.
New exciting algorithms such as Amazon's Dynamo (Werner Vogels is one of my favorite speakers and bloggers on distributed tech, if you are not familiar with him in this space you should be) database are showing in real world situations that distributed systems and distributed data are a reality.
In a lot of system designs because people are so familiar with relational data structures and systems we find Object Models that look like a relational database design. When asked why it looks like this the answers are fairly consistently things like "this is how the database stores it, for speed we need to do the same" or "it just made sense when we pulled the DBA in to help us with the model."
Objects are not relational! They are objects. Then when you get into full structures of objects or object trees there are relationships but it is not the same as a relational database. Especially as we start to build and mature distributed system algorithms it doesn't make sense to use a centralized data store. If the data can be broken up, distributed and stored with the algorithms that will use it performance will improve.
In fact I would argue that the elusive SLA of a system response can begin to be discussed if we can tie the data to the processing. Granted there are new complexities in this model for synchronization, segmentation and consistency but there are ways to solve them. Similarly consistent access to the same servers is also possible.
What other great examples of distributed computing and distributed data storage have you seen?
Tuesday, June 26, 2007
Every integration point created WILL fail and Scalability
I have two things for today's post.
1. I received a note this AM from my "FYI- Jana RSS feed" (A good friend that works with me that sends me email) that I had to pass on to everyone. Read the article because the quote below is SOOOOOOO true.
“One of my recurring themes in "Release It"
is that every call to another system, without exception, will someday try to
kill your application. It usually comes from behavior outside the specification.
When that happens, you must be able to peel back the layers of abstraction, tear
apart the constructed fictions of "concurrent users", "sessions", and even
"connections", and get at what's really happening…”Agile, Architecture and
the 5am Production Problem - by Michael
Nygard
2. This weekend I attended the Seattle Conference on Scalability. It was sponsored by Google and it was GREAT. Over the next few weeks I will digest what I heard and saw and report it back via blog. At the moment though I have to admit my dueling day jobs are getting the better of me so I haven't had time to couch my ideas into pithy post prose. I promise it is coming though. I will however let you know that Google recorded the presentations and they should be available by the end of the week on YouTube.
Sunday, May 27, 2007
Observations from my Romp with Ruby
My experience with Ruby on Rails was a positive one, or to quote the Rails site, it was "Web development that doesn't hurt". I am still by NO stretch a ruby or rails expert but I did learn a lot and I was impressed with the abilities in such a relatively young language. I still believe that no language will ever solve all of the worlds programming ills. But they will continue to get to higher and higher levels of abstraction and developer productivity.
Ruby is essentially an object oriented scripting language. Rails is a separate library set that is focused on productivity enhancements and tools to really accellerate web development. Rails uses the Model View Controller pattern for development and a strong implementation of the Domain pattern for data persistence. You could very easily build a fairly complex app all without any knowledge of the database. (Now there are obviously other concerns with that but we won't dig into that in this post.)
CSS and stylesheets can be automatically generated along with full basic UIs with a tool called scaffolding.. (It's not that pretty but gives something upon which to build and actually works great for validating that you have the right model). For many of the common advanced UI actions there are simple sets of commands as well with built in Javascript AJAX - scriptaculous commands for visual effects, drag and drop, dynamic lists and more. Everyone knows you have to have AJAX to be cool now adays.
If you want to poke at Ruby and learn a bit check out RadRails for development (it lacks intillisense type function at the moment which made me sad but everything else you could want works right out of the box) and InstantRails for a working environment. It's a great combo to get you going in just a few short downloads.
Every language will have some weak spots and Ruby is no exception. But it does work hard to answer many of the common problems in today's frameworks. My favorite (have to give a plug for the Engineering side of things) is things like fixtures and test automation built-ins to encourage test driven development. My biggest question is still, how does it scale with very high volumes?
Sunday, April 15, 2007
A quick snip on testing
I have recently had two different offline topics end up with the same point so I figured I would pass it on here as well. Testing and why it is important to test to failure and not just test. To stick to the intuitive reason, it is because then you know when your system will break and what will happen when it does. It is naive to believe that your system will be 100% problem free, it is a far more realistic expectation to accept that you just don't know when.
One of the fundamental race conditions of the universe is that as soon as you build a better, more idiot-proof system, the world will build a better idiot. Similarly in a system which is successful and experiencing growth, you will eventually hit a point that exposes problems that didn't show themselves under lesser pressure or lesser load. In aircraft they X-Ray and image things like propellers not because they want to ensure that it is perfect, but so that they know what quality it is. Under the right conditions a very tiny bubble can cause a propeller to literally fly apart. Better to know these things up front.
My Jana RSS feed recently (inside joke) recently came across with these two links, one for example, one for humor that are relevant to this topic. Check out this Ultimate Failure test... and fly with a little better feeling in your heart because you know that they care enough to do this. Then check out Monkey Testing at it's best. (I am trying very hard to resist the 1000 monkeys and Shakespeare sonnet reference... darn... I failed.)
In conclusion remember, it's ALWAYS better that you know where things will blow up than for someone else to tell you.
Friday, February 16, 2007
Not dead yet...
Rumors of my death have been greatly exaggerated. Ok. So maybe there weren't really rumors of death. I did hear from enough of my regular readers though that I wanted to pop a quick post out to declare that I am still alive just HORRIBLY busy at the moment. I normally write my blogs at night and post when I get a chance but of late I have been working regular work during that time. Rude isn't it?
Rather than simply declare that I am alive (and whine that I am busy) I will share a bit of scalability learning as well. I was recently sent a link to a wonderful article on the stability journey of MySpace.com (if you don't know what MySpace is you probabl live under an internet rock) in Baseline Magazine. It's a fairly quick read and it discusses many of the standard things you learn with scale. (Partitioning, Virtualization, Simplicity, Segmentation etc.)