Saturday, June 03, 2006

Part 1: Disruptive Innovation in the path from technology to brand - a maturity model | by Adrian Cockcroft | June 3rd, 2006

Products aim to fill a need in a market, products that are disruptive innovations also reshape the market and markets tend to evolve in a series of discontinuous steps as they mature. The phrase "crossing the chasm" has been used to describe these changes, and "early adopters" are the people who first move a market to a new phase.

In the next few posts I'm going to describe a generic maturity model that applies to many markets, and show how disruptive innovations may drive a market into a more mature phase. I got the initial idea of looking at markets in this way from Dave Nocera of Innovativ in a presentation he gave at SUPerG in early 2004. He used video as an example, with the move from VCR to Video rental to Online. I have extended that example, and come up with a generic maturity model based on it, which I also apply to the Telco industry.

Monday, May 15, 2006

See you at the developers conference? | by Adrian Cockcroft | May 15th, 2006

The combined eBay, PayPal and Skype developer conference is coming up, June 10-12 in Las Vegas. I missed the event last year, but I will be staffing it this year! A few of us are being let out of the mysterious eBay Research Labs for the occasion. They told us to get to work on the future of e-commerce, and have kept us locked up for months, shipping in occasional supplies of Starbucks and fresh interns. I've been writing serious amounts of code for the first time in years, if fact I'm too busy writing Java to have time to go to JavaOne this week.

The conference is supposed to illuminate questions such as:

- What will the next technology revolution be

- How will it impact commerce and communications on the web

- And what opportunities will it provide for developers and technology innovators

- How will the Long Tail Theory play out

- Web 2.0 and how to build revenue streams

Its become a common joke to keep incrementing this: Web 2.1, Web 3.0 etc. but personally I think the most interesting developments aren't even Web based. e.g. Skype isn't a Web application, it defines its own virtual private peer to peer fabric that overlays the Internet.

See you in Vegas!

Wednesday, April 19, 2006

Blogging Tools | by Adrian Cockcroft | April 20th, 2006

I've been using blogger for the last 18 months, it was an easy way to get started but now I don't see some of the features I want. The basic service has changed very little in that time, so it doesn't seem to be getting much investment and development.

The three missing features I see in other blogs are tags, a blogroll, and posting categories.

I want to have an easy way to add a series of tags to each blog entry, without having to create custom html. I did it once the hard way and don't usually bother.

I'd like a blogroll so that people can see which blogs I think are worth reading, but I don't want to edit my html template to get one, I want to import OPML or have a table to edit.

I'd like to be able to separate categories so that I can label rants like this separately from technical info on capacity planning, thoughts on the industry, personal stuff.

I like the web based blogger service, I can post from anywhere using any device (I've posted to blogger from Linux, Solaris, Windows, Mac and Treo/palmOS). I don't want to host my own blog or have to install a blogging tool.

I use bloglines as an aggregator to read blogs, I could also use bloglines to host my own blog, since it does seem to have some of these features, and it would make referring easier.

What other options are out there, is there a slightly better blogger competitor that I should check out? Is there a way to migrate existing entries to a new blog? Comments requested...

Cheers Adrian

Comparing Smart Mobile Phones | by Adrian Cockcroft | April 19th, 2006

Its been a while since I last posted, mostly due to a long vacation. We stayed with friends in New York for a few days, spent 10 days on Bermuda (very nice and relaxing) and spent a few days in New York again on the way home.

In the last week or so I have been trying out a new phone. I got a Nokia 6682 which runs the Symbian S60 operating system and which has fairly good third party support for applications that use Java, Flash and Opera. My history with phones started out with Nokia for many years, and then I switched to the Treo line. I've had a Treo 270, 600 and currently have a 650. Going back to Nokia was is some ways familiar, the user interface has some similarities to the older days, but overall I miss my Treo and I'm going to switch back. I thought it might be interesting to discuss the differences and what I think a state of the art smartphone should be able to do for me.

The Nokia 6682 has a decent spec, large color screen, 1.3Mpixel camera and a 64MB removable flash card included. The spec page even says that you can "Bid on eBay on the go" but that is available to any phone that can browse to wap.ebay.com. With a normal Cingular GSM service its not using a 3G high speed network, so data network access is similar in speed to the Treo 650.

My main problem is that I've been spoilt by the Treo's touch screen and keyboard. When using the Nokia, at first I was poking at the screen in vain trying to select things. The screen is fairly high resolution, but its an eye test in that the text size is too small for many features, and the colors available in the default set of themes have poor contrast. The Nokia is actually much harder to read. For text entry I'm actually used to using Nokia's predictive text feature but it is still extremely painful to enter a text message or URL.

The Treo' browser (Blazer) is easy to use but is not well supported in terms of javascript and many sites don't recognize it properly. On the Nokia there is a built in browser (called Web) and Opera 7 is included, with a free upgrade to Opera 8.5. I found the "Web" browser OK to use, but Opera was very annoying and unintuitive. With the Treo I can quickly get on the web to look something up, it just takes too long on the Nokia, and with Opera I found the navigation commands to be confusing and awkward. I tried to make use of the javascript support in Opera, but it didn't work for Google maps, despite Google claiming that Opera 8 is supported.

I also had problems with the Nokia after browsing the web. The phone keeps applications running in the background and tends to run out of memory at awkward moments. I tried to use the camera, but at the point of taking a picture it failed with a lack of memory. I had to bring up the web browser and explicitly exit it, meanwhile the photo opportunity had gone. The Nokia's camera is higher resolution and has continuous zoom that works for video as well as pictures. For the Treo, you have just 1x and 2x zoom settings and a 640x480 resolution. Its not really enough to snap pictures of whiteboard scribbles clearly. The Nokia has a sliding cover for the phone, which activates the camera when opened, but its too easy to open by accident when getting the phone out.

Other phones I've seen recently include the Verizon LG VX9800 which is a very fat clamshell with a nice big keyboard, 3G networking and a real eye-test of a small hi-res screen. A friend got one but is taking it back, its more suited to gaming and entertainment than business. My son has a Motorola SLVR L7 with iTunes and seems happy with it. It looks cool, fits his interests but he doesn't try to use the web from his phone.

Some co-workers have Windows mobile phones, I haven't tried to use them myself, but I've heard a mixture of good and bad comments, opinion seems very polarized as love it or won't touch it..

So in summary, the things I can't do without on a phone are a touch screen, full keyboard and a large (not just high resolution) display with fonts and icons that can be read easily. The things I don't like about the Treo 650 are its lack of support for Opera 8 (which may be more usable with a keyboard and touch screen), Javascript and Flash.

My favourite applications on the Treo are the Chatter email client, Planetarium for identifying stars and planets, and Solitaire for mindless time wasting - which would be a pain to play without the touch screen. I also find that the mobile version of bloglines works well with the Treo's browser, so I can keep up with my feeds.

Tuesday, March 14, 2006

How to finish writing a book | by Adrian Cockcroft | 15th March 2006

I've written four books, and several years ago I developed "Cockcroft's law of book writing". This states that a book will grow in size as you write it, and that the number of pages left to write will increase as you write. This seems counter-intuitive, but it has been confirmed many times in practice. I hope this posting provides some useful advice for writers, and helps people finish what they have started.

To make a concrete example, let's say you decide to write a book and you come up with an outline that adds up to 200 pages. You start work and write 50 pages, then, when you revisit your outline to update the page count estimates, you find that they now add up to 300 pages. You wrote more than you expected to cover each subject, and discovered more subjects that needed to be discussed. The essential problem here is that there are now 300-50 = 250 pages left to go. Before you started you only had 200 pages left to go.

This problem is recursive, if you write another 50 pages you will find that you have now written the first 100 pages of a 400 page book, and you now have 300 pages left to write. This explains why there are so many people who have written part of a book, but never finished it.

The aproach I took in writing my later books was to maintain a spreadsheet that tracks the pages left to write or edit, update it very regularly, and generate a plot with a trend line from the data. You can then see when (or if) you will finish the book. In order to get the trend line to target a specific delivery date, you have to force the number of pages left to go down. You do this by writing pages that you promise never to edit again, and by deleting whole sections and chapters. I deleted three entire chapters from one of my books to get it finished.

Another problem you can run into is that the content you wrote at the start of the process is less well written than later content, so you think you have finished, re-read parts of the book that were finished ages ago, and discover that it needs a complete rewrite.

I often get asked if I will update my Sun Performance and Tuning book, and I don't intend to do a third edition. This is mainly because I'm interested in other things, and I'm no longer up to date with the subjects I would need to cover. I have sketched out a possible book on capacity planning with free tools, and the trend line on that book is nice and flat. I haven't really started writing it, and so it hasn't started getting bigger yet....

My good friends Jim Mauro and Richard McDougall are closing in on the end point for Solaris Internals 2nd Edition. I've been looking forward to it for a while, and its going to be a monster book, covering how Solaris 10 really works, lots of DTrace based examples, and is going to be the essential companion for anyone looking at Open Solaris.

Saturday, March 11, 2006

Strange Contextual Ads | by Adrian Cockcroft | 11th March 2006

The eBay contextual ads seems to be fixated on the word "margin" which is not present in the content of my blog. However the CSS template that is used to host this blog contains the word margin over and over again. I think Alex needs to do a better job of filtering out formatting words before he does his context analysis....

Thursday, March 09, 2006

Contextual eBay adverts with ctxbay | by Adrian Cockcroft | 9th March 2006

The winners in the eBay developer contest were announced at ETech, one was Alex Stankovic, who has developed a contextual advertising system for eBay that works just like Google Adsense. I just changed my advert bar for this blog to use his system at ctxbay.

At the ctxbay site you find a link to sign up with eBay for the affiliate program (via Commission Junction) this was an easy fast signup, and gets you an affiliate id number. You then create an account at ctxbay, and enter your affiliate number, which they will then use to call back to Comission Junction and make sure you get paid.

The rest of the setup is similar to adsense, however you do need to login to ctxbay with your new account, and it didn't do this automatically for me. ctxbay generates a selection of common ad frame formats using javascript that can be slotted into your site template. I found one identical to my adsense format, and swapped out the code. My first attempt didn't work because I was not logged in, and the id field in the javascript was empty. After I logged in I got a fairly long string that keys my ad to ctxbay.

I setup adsense in order to understand it better, and going forward I'll see if ctxbay can generate any sensible eBay items out of the keywords in my blog.
,

Wednesday, March 08, 2006

Etech Tuesday On Rails | by Adrian Cockcroft | 8th March 2006

A long day with lots of interesting talks, and I got to chat with several new people and also to try out my FLORWAX pitch. "Its the equivalent of AJAX but for Wireless" is my instant summary. To get this out of the way, the general reaction is that the name gets a chuckle (not too many groans yet), that there really is a big problem in wireless platform fragmentation, and that no-one seems to know of any other initiatives that have picked on this as a problem to solve. I think most people look at wireless, see this problem, and give up, as its too hard to make something work. Since I'm interested in a longer term perspective than most people, its seems fair game to try and provoke a discussion on what a sensible core set of wireless platform technologies would look like.

As Jesse James Garrett said in the tutorial yesterday, the key elements of AJAX are that it uses a common standard bundle of browser based technologies and that it is asynchronous, so you don't have to click-and-wait......click-and-wait......
If we apply these principles to Wireless, we need to define a bundle of standard technologies (I suggest Flash Lite 2.0, Ruby on Rails, XML web services - FLORWAX = FlashLiteOnRailsWirelessAsynchronousXml, but the actual bundle doesn't matter as long as a common set emerges). However the asyncronous problem is far worse in wireless than in desktop applications, we really need to have wireless apps that talk to the backend and update the screen without the click-and-wait-for-ages mode that is the norm.

At the end of the day I attended the Ruby on Rails BoF. I have heard good things about RoR but haven't used it. I think they converted me, and I took the opportunity to mention FLORWAX to the group. It does seem like the right technology fit.

The conference itself started with Ray Ozzie showing how to do cut and paste on the web. It seems so trivial, why hadn't been done before? A very useful way to make web apps behave more like regular apps. We then had a very cool hardware demo by Jeff Han, he has a touch screen that can see all his fingers separately and has created a very nice new set of user interaction paradigms.

Amazon has created a way to harness real people to do the stuff that AI can't do. Its called the Mechanical Turk, and its another simple idea with quite profound and wide reaching uses. Dick Hardt from Sxip gave an interesting talk on indentity, but the way he presented it with one word per slide and rapid fire transitions reminded me of Steve Colbert presenting his "The Word" section on The Colbert Report. I enjoyed it but I don't remember much of the content.

Next we has a talk from Felix Miller of last.fm on how they collect the metadata on what you are listening to and use it to help you find new music, I've been playing around with Pandora and training it to play the music I like, and I think I'll have to have a go at last.fm as well. I have eclectic tastes, and its hard to keep the recommendations from veering back to the mainstream in Pandora.

After the break, there were several presentations that didn't grab my attention or told me things that seemed obvious to me. The highlight was a presentation on Second Life that was presented using a billboard in the virtual world and lots of interactive explanations of how it all works. Fascinating, but I don't have enough time to play as much as I'd like in the real world....

, ,

Tagging Blogger and Technorati | Adrian Cockcroft | 8th March 2006

Blogger doesn't provide an integrated way to tag my postings (as far as I can tell) but at ETech we were advised to tag our blog entries with Etech and Etech06 so I tried to figure this out during one of the less interesting talks (yes I know I should have figured this out ages ago). Despite a slow and intermittent wireless internet connection I found that I could use Technorati to do this by embedding some html in my blog. The frustrating problem I ran into was that Technorati seemed to be completely overloaded and was largely unresponsive. When I did finally get my blog entry tagged I went to Technorati and searched for Etech, and after a long wait it came up with no results at all... I figured I had done something wrong, but late at night the load dropped off and the site now does actually find Etech tags, including my own one. I celebrated by adding "florwax" as a tag, so we will see how that goes...

Another site that seems a victim of its success is Myspace, their music delivery service has become overloaded, so its hit and miss whether you can get any songs to play at the moment.

,

Monday, March 06, 2006

Thoughts from ETech - Tutorial Day - FLORWAX? | by Adrian Cockcroft | 7th March 2006

I'm at the O'ReillyEmerging Technology conference in San Diego, today was "Tutorial Day" and I decided to attend "Designing the next generation of Web Applications" in the morning and "Next Generation Flash Development with Flex" in the afternoon.

The morning talk was very nicely presented, Jesse James Garrett coined the term AJAX amongst other things, and Jeff Veen worked on Hotwire, Blogger and Measuremap, they provided a structured set of best practices for designing web applications with lots of great examples and anecdotes.

AJAX acts as a convergence point for browser based applications because almost all current browsers support the same set of technologies and there is a highly functional lowest common denominator. There is now also a large body of applications that provide the intertia or value that constrains the browser writers from diverging with incompatible functionality. This is the same effect that occurred in the PC marketplace, when MSDOS and Windows developed enough application value that neither Intel nor Microsoft could diverge in an incompatible manner. The collective self interest of the end user reaches a tipping point that blocks radical innovation, and slow incremental evolution takes over.

The interesting area for me is how this maps to the mobile/wireless space. There is no AJAX for wireless applications, the market is huge, but the platform diversity is also huge and is growing. One estimate I heard was that there are 1000 separate platforms to target and that this number is growing, not shrinking.

So what we need, is the equivalent of AJAX for Wireless, something like Flash Lite On Rails Wireless Asynchronous XML - FLORWAX - which also has a household cleaning connotation :-)
If enough people standardize on a common set of technologies, then the handset vendors will start to build to a common profile as well, and we could end up with a decently functional lowest common denominator.
, ,

Friday, February 24, 2006

Conferences and Innovation

I just signed up for the O'Reilly Emerging Technology event in San Diego next month - http://conferences.oreillynet.com/etech/

I've also written a paper for a workshop in the IEEE Joint Conference on E-Commerce Technology (CEC'06) and Enterprise Computing, E-Commerce and E-Services (EEE'06) http://linux.ece.uci.edu/cec06/ - but the conference name is so long that I can't remember it very well in conversation. This conference also includes the 2nd International Workshop on Business Service Networks (BSN '06) and the 2nd International Workshop on Service oriented Solutions for Cooperative Organizations (SoS4CO '06). Its all sounds very interesting, its in June in San Francisco, and needs a snappier name...

Last December I attended the Fortune Innovation Forum in New York. It was very nicely put together and in effect it validated the approach we were already taking. It seems that most of the attendees were trying to work towards a culture, process and tools for fostering innovation that seemed similar to our own setup. eBay and PayPal were used as examples several times.

We used a few simple techniques last year to kickstart our own innovation program. One method I borrowed from other events is the "Poster Lunch". Get a room near the company cafe, provide flip chart sized pads and pens, email everyone to tell them about it and put up signs in the Cafe to invite them in on the day. Anyone can put anything they like on a poster, stick it up and collect comments on it in person. One thing we found was that there were several posters suggesting eBay site features that already existed or were in development. One suggestion in particular was getting lots of support and comments until someone wrote on it "LTS thursday!", meaning it would be Live To Site and be launched two days later. We also gave attendees voting stickers so that they could indicate their favourite posters.

To drill down on the best ideas we also setup a regular open-to-all meeting where we could discuss the concepts and route them to the appropriate expert or business owner. The most far-sighted ideas get routed to become candidates for research labs projects, and the people who had the ideas get to develop them further.

To support the collection of ideas, we created a Wiki. This is nice because it is free format, and supports comments and discussion, with very low initial barrier to entering an idea. The problems came when there were several hundred ideas in the Wiki, it became hard to maintain. A more specialized pre-concept tool that feeds into our standard development process is a better solution for incremental innovations, and the Wiki works better for more radical ideas.

To really get a dose of innovative ideas, last year I attended a seminar on Complex Adaptive Systems by the Santa Fe Institute. It was a real eye-opener, they are pushing the boundaries of multi-disciplinary research, e.g. forming teams with Physicists, Biologists and Economists to derive the rules of scaling and organization of living things, from the smallest mammal to the largest city. Since eBay, PayPal and Skype are social networks, (their value comes from connections within their communities) they behave in some ways like cities, and follow similar kinds of scaling rules.

Saturday, February 18, 2006

Changing gears

I started this blog in the summer of 2004 when I had finished at Sun and not yet started at eBay. After 16 years at Sun this was a big move. I knew people at eBay from the time in 1999 when they had a big outage and many Sun people got involved in helping them get up and running again. My thinking that summer was that web services platforms were where the real innovation was taking place, and I see eBay and PayPal as the leading transactional web services platforms.

My first year at eBay was in the Operations Architecture group, where I was working on figuring out new platforms and upgrades, and helping with capacity planning tools and processes. I also figured out a lot about how eBay and PayPal really work, and the challenges of scaling a rapidly growing and changing high availability transactional platform to a size that is beyond most people's comprehension. I had some entertaining meetings with hopeful vendors who would come in with solutions to common industry problems (e.g. low utilization) that eBay doesn't have, and their largest existing deployment would be an order of magnitude too small to be useful. After describing a bit about how eBay works, they would get big eyes, admit that their product wasn't appropriate, and wander off to look for more normal customers... A lot of what eBay does is built internally because the generic products don't scale and we can build what we need ourselves for less.

In the summer of 2005 I moved internally to help form eBay Research Labs. Since then we have hired some very experienced researchers and are becoming the focal point for innovation within eBay. This was another opportunity for me to change gears and greatly increase the scope of my work. Part of my role is to continue to research new platforms and technologies for the datacenter operations, and I've been joined in this work by my friend Paul Strong. Paul was in the N1 group at Sun, and is also the drummer for Fractal. Paul and I were both involved in the Enterprise Grid Alliance, he ended up as chair of the Technical Steering Committee, and edited the EGA's Grid Reference Architecture. He's now working on how to enhance the automation of eBay's datacenters.

The other cool thing that came my way in 2005 was eBay's purchase of Skype. Its not just a VOIP tool, its a huge and fast growing community (something eBay understands very well) and an extremely innovative development platform. The Skype API is a fun place to do innovative research, and the Skype network has between 3 and 5 million active nodes at any point in time (up by a million in three months). I've been interested in the telecom market ever since I was one of the Sun Systems Engineers working with British Telecom in the early 1990's. Now I get to play with the future of telecom in the form of Skype, and I'm also very interested in mobile/wireless applications.

In another sense I am changing gears with this post. I've changed the title and description, and it is now also being included in the Best of eBay Blogs site. I've been encouraged by the example of other bloggers at that site to discuss a bit more openly what I get up to, but if you ask me what I'm really working on, all I can say is "The future of e-commerce".

Cheers Adrian

p.s. I just tried to spell-check this posting, and the built-in spell checker at blogger.com decided that the first error was the word "blog", which I find highly amusing, so I gave up and any spelling errors in the above are my fault.

Monday, January 30, 2006

Interesting hardware for database servers

I've been too occupied on other things to keep posting regularly in the last month. The good news is that I'm learning a lot about some new areas.

So what is new in hardware? I think there are some interesting trends in server hardware for running databases. The cost base of a mid-sized Solaris/Oracle server with 32-64GB of RAM is dropping fast. A very common platform in this space has been the 8-way UltraSPARC III based V880, moving to the 12-way V1280 and currently the E2900 (a 24 core V1280 chassis with UltraSPARC IV) over the last few years. Prices vary by configuration, and newer systems give you more performance per $, but are of the order of magnitude of $100K (plus the disk subsystem and software licenses - but thats another topic).

The two new entrants in this space are Niagara based systems and Opteron based systems, each has its strengths. When loaded up with RAM the costs are largely dominated by the price of RAM rather than the CPU itself, however both these systems use commonly available DIMMs, rather than the more specialized and expensive memory of the older generation systems.

Niagara has everything on one chip, 8 cores and 32 execution threads. The cool thing about this for database is not the low power consumption touted by Sun (which is dwarfed by the disk subsystem for a database application) , but is that any inter-thread locking will be blindingly fast since the signals do not have to go off-chip. Badly behaved applications that are sensitive to high memory latency (the kind that don't scale well on physically bigger systems) will run relatively well. However the cores themselves are not particularly fast and are atrociously slow for anything that does floating point, so single stream performance is not a strength. With 2GB DIMMs you can get 32GB RAM connected to a single Niagara chip, this should move to 64GB using 4GB DIMMs eventually. The performance of a Niagara seems to be a bit better than a V1280, but the cost is much much lower. Software support for SPARC Solaris 10 doesn't seem to be an issue at this point. Most things are supported and the system is compatible with earlier releases of SPARC/Solaris products.

The common Opteron systems are two socket/four core with a maximum of 16GB with 2GB DIMMs. There are some four socket and eight socket systems available from several vendors (including Sun), with 32-64GB, moving to 128GB with 4GB DIMMs. The Opteron seems to have performance per GHz in the same ballpark as UltraSPARC systems, so 8 cores at 2.4GHz would be between the performance of a V1280 and E2900. The 32GB 8-Core Opteron systems are in the same order of magnitude for performance and price as a 32GB Niagara, but far faster for single stream work and floating point, and relatively slower for lock intensive workloads where the signals have to move between the Opteron chips. On Opteron the software situation is a little different, there are three possible operating systems - Solaris 10, Linux and Windows 64. Solaris support isn't as good as it is on SPARC, for example Oracle 10g is the only option, the earlier releases of Oracle don't seem to be available. Linux probably has the widest choice for support, but 64bit Linux on larger systems doesn't seem to scale as well as Solaris 10 in my experience. Linux tends to be more efficient than Solaris on 32bit systems (your milage will vary, it depends greatly on what features of the OS your workload hits hard). I don't know anything about Windows 64, but I expect these large Opteron systems will be good SQLserver platforms.

Thats what the landscape looks like to me as we go into 2006. I hope to be doing some testing later this year to compare all the options, including Intel's next generation servers, to get my performance and price comparisons to be more precise than the general comments above. I'd be interested to swap experiences with other people moving in this direction.

Wednesday, December 21, 2005

CMG05 trip comments and "utilization is useless..."

I have a good time at the Computer Measurement Group meeting in Orlando recently. Mario Jauvin and I put together a tutorial on "Capacity Planning and Performance Monitoring with Free Tools" that was well attended, although relatively few attendees seem to be using free tools, mostly due to policy and support issues. We also attended James Holtman's workshop on using the free statistics package 'R' for performance data analysis and plotting.

The main new theme at this year's conference seemed to be CPU virtualization. Many people using VMware, Zen, Solaris containers and other virtualization facilities are finding that their measurements of CPU utilization don't make sense any more. Both BMC and Teamquest are working on building some support for virtualization concepts into their tools.

My observation is that utilization is useless as a metric and should be abandoned. It has been useless in virtualized disk subsystems for some time, and is now useless for CPU measurement as well. There used to be a clear relationship between response time and utilization, but systems are now so complex that those relationships no longer hold. Instead, you need to directly measure response time and relate it to throughput. Utilization is properly defined as busy time as a proportion of elapsed time. The replacement for utilization is headroom which is defined as the unused proportion of the maximum possible throughput. Dave Fisk calls this Capability Utilization.

I'm thinking of writing a paper, maybe for next year's CMG on this topic....

Happy holidays everyone
Cheers Adrian

Monday, October 31, 2005

Help! I've lost my memory! Updated Sunworld Column

Originally published in Unix Insider 10/1/95
Stripped of adverts, url references fixed and comments added in red type to bring it up to date ten years later.

Dear Adrian, After a reboot I saw that most of my computer's memory was
free, but when I launched my application it used up almost all the
memory. When I stopped the application the memory didn't come back!
Take a look at my vmstat
output:

% vmstat 5
procs memory page disk faults cpu
r b w swap free re mf pi po fr de sr s0 s1 s2 s3 in sy cs us sy id

This is before the program starts:

0 0 0 330252 80708   0   2  0  0  0  0  0  0  0  0  1   18  107  113  0  1 99
0 0 0 330252 80708 0 0 0 0 0 0 0 0 0 0 0 14 87 78 0 0 99

I start the program and it runs like this for a while:

0 0 0 314204  8824   0   0  0  0  0  0  0  0  0  0  0  414  132   79 24  1 74
0 0 0 314204 8824 0 0 0 0 0 0 0 0 0 0 0 411 99 66 25 1 74

I stop it, then almost all the swap space comes back, but the free memory does not:

0 0 0 326776 21260   0   3  0  0  0  0  0  0  1  0  0  420  116   82  4  2 95
0 0 0 329924 24396 0 0 0 0 0 0 0 0 0 0 0 414 82 77 0 0 100
0 0 0 329924 24396 0 0 0 0 0 0 0 0 2 0 1 430 90 84 0 1 99

I checked that there were no application processes running. It looks like a huge memory leak in the operating system. How can I get my memory back?
--RAMless in Ripon

Update

This remains one of the most frequently asked questions of all time. The
original answer is still true for many Unix variants. However while
writing his book on Solaris Internals, Richard McDougall worked out
how to fix Solaris to make it work better, and to make this aparrent
problem go away. The result was one of the most significant
performance improvements in Solaris 8, but the first edition of his
book was written before Solaris 8 came out, so doesn't describe the
fix!

The short answer

Launch your application again. Notice that it starts up more quickly than it did the first time, and with less disk activity. The application code and its data files are still
in memory, even though they are not active. The memory they occupy is
not "free." If you restart the same application it finds
the pages that are already in memory. The pages are attached to the
inode cache entries for the files. If you start a different
application, and there is insufficient free memory, the kernel will
scan for pages that have not been touched for a long time, and "free"
them. Once you quit the first application, the memory it occupies is
not being touched, so it will be freed quickly for use by other
applications.

In 1988, Sun introduced this feature in SunOS 4.0. It still applies to
all versions of Solaris 1 and 2. The kernel is trying to avoid disk
reads by caching as many files as possible in memory. Attaching to a
page in memory is around 1,000 times faster than reading it in from
disk. The kernel figures that you paid good money for all of that
RAM, so it will try to make good use of it by retaining files you
might need.

Since Solaris 8, the memory in the file cache is actually also on the free
list, so you do see vmstat free memory reduce when you quit a
program. You also should expect large amounts of file I/O to cause
high scan rates on older Solaris releases, and for there to be no
scanning at all on Solaris 8 systems. If Solaris 8 scans at all, then
it has truly run out of memory and is overloaded.

By contrast, Memory leaks appear as a shortage of swap space after the
misbehaving program runs for a while. You will probably find a
process that has a larger than expected size. You should restart the
program to free up the swap space, and check it with a debugger that
offers a leak-finding feature (run it with the libumem version of the malloc library that instruments memory leaks).

The long (and technical) answer

To understand how Sun's operating systems handle memory, I will explain how the inode cache works, how the buffer cache fits into the picture, and how the life
cycle of a typical page evolves as the system uses it for several
different purposes.

The inode cache and file data caching

Whenever you access a file, the kernel needs to know the size, the access permissions,
the date stamps and the locations of the data blocks on disk.
Traditionally, this information is known as the inode for the file.
There are many filesystem types. For simplicity I will assume we are
only interested in the Unix filesystem (UFS) on a local disk. Each
filesystem type has its own inode cache.

The filesystem stores inodes on the disk; the inode must be read into
memory whenever an operation is performed on an entity in the
filesystem. The number of inodes read per second is reported as
iget/s by the sar
-a
command. The inode read from disk is cached in case
it is needed again, and the number of inodes that the system will
cache is influenced by a kernel parameter called ufs_ninode.
The kernel keeps inodes on a linked list, rather than in a fixed-size
table.

As I mention each command I will show you what the output looks like. In
my case I'm collecting sar
data automatically using cron.
sar, which defaults to
reading the stored data for today. If you have no stored data,
specify a time interval and sar
will show you current activity.

% sar -a

SunOS hostname 5.4 Generic_101945-32 sun4c 09/18/95

00:00:01 iget/s namei/s dirbk/s
01:00:01 4 6 0

All reads or writes to UFS files occur by paging from the filesystem. All
pages that are part of the file and are in memory will be attached to
the inode cache entry for that file. When a file is not in use, its
data is cached in memory, using an inactive inode cache entry. When
the kernel reuses an inactive inode cache entry that has pages
attached, it puts the pages on the free list; this case is shown by
sar -g as %ufs_ipf.
This number is the percentage of UFS inodes that were overwritten in
the inode cache by iget and
that had reusable pages associated with them. The kernel flushes the
pages, and updates on disk any modified pages. Thus, this %ufs_ipf
number is the percentage of igets with page flushes. Any non-zero
values of %ufs_ipf reported by sar -g
indicate that the inode cache is too small for the current workload.

% sar -g

SunOS hostname 5.4 Generic_101945-32 sun4c 09/18/95

00:00:01 pgout/s ppgout/s pgfree/s pgscan/s %ufs_ipf
01:00:01 0.02 0.02 0.08 0.12 0.00

For SunOS 4 and releases up to Solaris 2.3, the number of inodes that the
kernel will keep in the inode cache is set by the kernel variable
ufs_ninode. To simplify: When a file is opened, an inactive
inode will be reused from the cache if the cache is full; when an
inode becomes inactive, it is discarded if the cache is over-full. If
the cache limit has not been reached then an inactive inode is placed
at the back of the reuse list and invalid inodes (inodes for files
that longer exist) are placed at the front for immediate reuse. It is
entirely possible for the number of open files in the system to cause
the number of active inodes to exceed ufs_ninode; raising
ufs_ninode allows more inactive inodes to be cached in case
they are needed again.

Solaris 2.4 uses a more clever inode cache algorithm. The kernel maintains a
reuse list of blank inodes for instant use. The number of active
inodes is no longer constrained, and the number of idle inodes
(inactive but cached in case they are needed again) is kept between
ufs_ninode and 75 percent of ufs_ninode by a new
kernel thread that scavenges the inodes to free them and maintains
entries on the reuse list. If you use sar
-v
to look at the inode cache, you may see a larger
number of existing inodes than the reported "size."

% sar -v

SunOS hostname 5.4 Generic_101945-32 sun4c 09/18/95

00:00:01 proc-sz ov inod-sz ov file-sz ov lock-sz
01:00:01 66/506 0 2108/2108 0 353/353 0 0/0

Buffer cache

The buffer cache is used to cache filesystem
data in SVR3 and BSD Unix. In SunOS 4, generic SVR4, and Solaris 2,
it is used to cache inode, indirect block, and cylinder group blocks
only. Although this change was introduced in 1988, many people still
incorrectly think the buffer cache is used to hold file data. Inodes
are read from disk to the buffer cache in 8-kilobyte blocks, then the
individual inodes are read from the buffer cache into the inode
cache.

Life cycle of a typical physical memory page

This section provides additional insight into the way memory is used. The sequence
described is an example of some common uses of pages; many other
possibilities exist.

1. Initialization -- A page is born
When the system boots, it forms all free memory into pages, and allocates
a kernel data structure to hold the state of every page in the
system.

2. Free -- An untouched virgin page
All the memory is put onto the free list to start with. At this stage the
content of the page is undefined.

3. ZFOD -- Joining an uninitialized data segment
When a program accesses data that is preset to zero for the very first
time, a minor page fault occurs and a Zero Fill On Demand (ZFOD)
operation takes place. The page is taken from the free list,
block-cleared to contain all zeroes, and added to the list of
anonymous pages for the uninitialized data segment. The program then
reads and writes data to the page.

4. Scanned -- The pagedaemon awakes
When the free list gets below a certain size, the pagedaemon starts to
look for memory pages to steal from processes. It looks at all pages
in physical memory order; when it gets to the page, the page is
synchronized with the memory management unit (MMU) and a reference
bit is cleared.

5. Waiting -- Is the program really using this page right now?
There is a delay that varies depending upon how quickly the pagedaemon
scans through memory. If the program references the page during this
period, the MMU reference bit is set.

6. Pageout Time -- Saving the contents
The pageout daemon returns and checks the MMU reference bit to find that
the program has not used the page so it can be stolen for reuse. The
pagedaemon checks to see if anything had been written to the page;
if it contains no data, a page-out occurs. The page is moved to the
pageout queue and marked as I/O pending. The swapfs code clusters
the page together with other pages on the queue and writes the
cluster to the swap space. The page is then free and is put on the
free list again. It remembers that it still contains the program
data.

7. Reclaim -- Give me back my page!
Belatedly, the program tries to read the page and takes a page fault. If the
page had been reused by someone else in the meantime, a major fault
would occur and the data would be read from the swap space into a
new page taken from the free list. In this case, the page is still
waiting to be reused, so a minor fault occurs, and the page is moved
back from the free list to the program's data segment.

8. Program Exit -- Free again
The program finishes running and exits. The data segments are private to
that particular instance of the program (unlike the shared-code
segments), so all the pages in the data segment are marked as
undefined and put onto the free list. This is the same state as Step
2
.

9. Page-in -- A shared code segment
A page fault occurs in the code segment of a window system shared library.
The page is taken off the free list, and a read from the filesystem
is scheduled to get the code. The process that caused the page fault
sleeps until the data arrives. The page is attached to the inode of
the file, and the segments reference the inode.

10. Attach -- A popular page
Another process using the same shared-library page faults in the same place.
It discovers that the page is already in memory and attaches to the
page, increasing its inode reference count by one.

11. COW -- Making a private copy
If one of the processes sharing the page tries to write to it, a
copy-on-write (COW) page fault occurs. Another page is grabbed from
the free list, and a copy of the original is made. This new page
becomes part of a privately mapped segment backed by anonymous
storage (swap space) so it can be changed, but the original page is
unchanged and can still be shared. Shared libraries contain jump
tables in the code that are patched, using COW as part of the
dynamic linking process.

12. File Cache -- Not free
The entire window system exits, and both processes go away. This time
the page stays in use, attached to the inode of the shared library
file. The inode is now inactive but will stay in the inode cache
until it is reused, and the pages act as a file cache in case the
user is about to restart the window system again. The
change made in Solaris 8 was that the file cache is the tail of the
free list, and any file cache page can be reused immediately for
something else without needing to be scanned first.

13. fsflush -- Flushed by the sync
Every 30 seconds all the pages in the system are examined in physical page
order to see which ones contain modified data and are attached to a
vnode. The details differ between SunOS 4 and Solaris 2, but
essentially any modified pages will be written back to the
filesystem, and the pages will be marked as clean.

This example sequence can continue from Step 4 or
Step 9 with minor variations. The fsflush
process occurs every 30 seconds by default for all pages, and
whenever the free list size drops below a certain value, the
pagedaemon scanner wakes up and reclaims some pages. A
recent change in Solaris 10, backported to Solaris 8 and 9 patches,
makes fsflush run much more efficiently on machines with very large
amounts of memory. However, if you see fsflush using an excessive
amount of CPU time you should increase “autoup” in /etc/system
from its default of 30s, and you will see fsflush usage reduce
proportionately.

Now you know

I have seen this missing-memory question
asked about once a month since 1988! Perhaps the manual page for
vmstat should include a better explanation of what the
values are measuring. This answer is based on some passages from my
book Sun Performance and Tuning. The book explains in detail how the
memory algorithms work and how to tune them. However the book doesn't cover the changes made in Solaris 8.

Tuesday, October 25, 2005

SunWorld Columns at ITworld

I found that the columns I wrote between 1995 and 1999 all seems to be online, but its hard to find, so I'm going to start by just listing everything I can find in its original form. BEWARE! Many of these articles were obsoleted by developments in later releases of Solaris, or refer to dead URL's so I'll work through them in subsequent blog entries and provide commentary and updates.

2001/03: Collected Short Questions and Answers
1999/08: What does 100 percent busy mean?
1999/07: Disk Error Detection
1999/03: Digging into the details of WorkShop 5.0
1999/02: SyMON and SE get upgraded
1998/12: Out and about at conferences
1998/11: What's new in Solaris 7?
1998/10: IOwait, what's the holdup?
1998/09: Do collision levels accurately tell you the real story about your Ethernet?
1998/08: Unlocking the kernel
1998/07: Clearing up swap space confusion
1998/06: How busy is the CPU, really?
1998/05: Processor partitioning
1998/04: Prying into processes and workloads
1998/03: Sizing up memory in Solaris
1998/02: Perfmeter unmasked
1998/01: SE Toolkit FAQ
1997/12: At last! The updated SE release has arrived
1997/11: Performance perplexities: Help! Where do I start?
1997/10: Learn to performance tune your Java programs
1997/09: Clarifying disk measurements and terminology
1997/08: How does Solaris 2.6 improve performance stats and Web performance?
1997/07: Dissecting proxy Web cache performance
1997/06: Analysis of TCP transfer characteristics for Web servers made easier
1997/05: The memory go round
1997/04: Craving more books on Solaris? Look no further
1997/03: How to optimize caching file accesses
1997/02: Increase system performance by maximizing your cache
1997/01: Design your cache to match your applications
1996/12: Tips for TCP/IP monitoring and tuning
1996/11: The right disk configurations for servers
1996/10: Solving the iostat disk mystery
1996/09: Unveiling vmstat's charms
1996/08: What's the best way to probe processes?
1996/07 Monitor your Web server in realtime, Part 2
1996/07: How can I optimize my programs for UltraSPARC?
1996/06: How do disks really work?
1996/05: How much RAM is enough?
1996/04: Monitor your Web server in real time, Part 1
1996/04: How does swap space work?
1996/03: Watching your Web server
1996/02: Which is better, static or dynamic linking?
1996/01: What are the tunable kernel parameters for Solaris 2?
1995/12: Performance Q&A Compendium
1995/11: When is it faster to have 64 bits?
1995/10 Help! I've lost my memory!

Sunday, October 02, 2005

How busy is your CPU, really?

Just in case you thought that you could compare your CPU utilization data across Solaris releases I have a few words of caution.

To start with there is the whole problem of interrupts, do they count as system time, or do they just make whatever they interrupted take longer?

Then there is the question of wait-for-io, and who is waiting for which io? This is a form of idle time that tends to confuse people as it doesn't really mean anything once you have more than one CPU.

There is one mechanism used to report the systemwide CPU utilization data. This data is reported in a per-cpu kstat data structure and is used by every tool that ever reports CPU usr, sys, wio, idle etc. Most tools sum the data over all the CPUs, mpstat gives you the per CPU data. The form of the data is a number of ticks of CPU time that accumulates, starting at zero at boot time. To measure CPU utilization over a time interval, you measure the difference in the number of ticks and divide by the time, and the tick rate. The tick rate is set by the clock interrupt, it defaults to 100Hz, but can be set to 1000Hz.

There are two mechanisms used to measure CPU usage, one is the old method of using the clock interrupt to see what is running every 100Hz. This is low resolution, and since the clock wakes up jobs, its possible for jobs to hide between the ticks, so its a statistically biased measure.
The other mechanism is microstate accounting, where every change of state is timed with a hires clock. This was only used for tracking CPU usage by processes, and needed special calls to get the data in releases up to Solaris 9.

So how do those nice numbers you get from vmstat or your favourite tool vary?

Solaris 8 and earlier releases: interrupt time does not get classified as system time

Solaris 9: Interrupt time is counted as system time.

Solaris 10: wait-for-io time will always be zero. The ratio of confusion to enlightenment was too high, so the entire concept has been removed from Solaris. This is a good thing.

Solaris 10 initial release: systemwide CPU time is now measured using microstates, this is far more accurate than ticks, but somehow the interrupt time ended up spread out rather than in system time.

This is unfortunate, but I'm told that an update of Solaris 10 will again classify interrupts as system time, and then it will all be just about as accurate as it can be.

You may be curious about the size of errors in these measurements, the answer is that "it depends". There is an SE toolkit script called cpuchk.se that looks at the process data and compares the ticks with the microstate data. However, I don't have any measures of interrupt time differences.

SE toolkit 3.4 on Solaris 10 Opteron Workaround

The current 3.4 build of SE is available from the sunfreeware site and supports Solaris 8, 9 and 10 for SPARC and x86. It also includes full source code of the interpreter for the first time.
However it only includes 32bit x86 support, and when run on a 64bit Solaris kernel on an Opteron system, it will fail to run. This is due to the return value from the isalist command returning amd64 as the first word, rather than pentium. You can get it going with a workaround by changing the startup script /opt/RICHPse/bin/se to match the string "*pentium*" rather than "pentium*". This lets the 32bit x86 binary run on the 64bit amd64 system, and most of the scripts will still work. Some that try to access the kernel directly viakvm will fail, but most scripts use kstat which doesn't need 64bit accesses.
The real fix is to use the source to compile an amd64 build....

Monday, September 12, 2005

SunWorld Offline

The monthly columns I wrote for SunWorld Online between 1995 and 1999 seem to have gone offline. At one point after SunWorld shut down I got permission from the editors at ITworld to take back the content, remove the adverts, and host it on the Sun site. Since I left Sun in 2004 the link has been closed down. I think there is still some useful content and historical interest in the columns, which seem to have been copied to a few other places already. I'm going to dig them out of my archives, and repost them here, along with comments and perhaps a few corrections and updates.

The original idea for the column came from a featured article on the Sun homepage in 1995 where I talked about the release of my first book, Sun Performance and Tuning - SPARC and Solaris. This generated an offer from the editors at SunWorld Online to do a monthly Performance Q&A column. I thought it worked out very well. I got paid for the column, which forced me to write something every month, and I based the columns on questions from the readers, or parts of the original book that needed to be updated, or parts of the second edition that I was writing during this period.

The monthly column seemed to have the effect of positioning me as the expert to quite a wide public, and I'm sure it helped me get promoted to Distinguished Engineer in 1999. I stopped writing the column due to a combination of being too busy, running out of topics to write about, and the introduction of the BluePrint program.

In effect, several engineers had created books and columns in their spare time, and we worked with management to make it an official program for publishing Sun's best practices on how to configure products and combinations of products. This made it our day job, which was easier in some ways, and we had support staff and more detailed reviews, but we didn't get paid royalties for doing the BluePrint columns or books. After a few years and several books (The very first BluePrint on Resource Management, and one on Capacity Planning) I moved on from the BluePrint group, but the program is still running and there is a large body of work there now.

There are over forty columns. I could post them en-masse unedited but I think they can already be found lurking in Google. It seems like more fun to try and remember what was going on at the time and make it an anecdotal reminiscence.

Monday, August 15, 2005

Solving storage tuning problems

I wrote a while ago about Dave Fisk's Ortera Atlas tool for storage analysis. I recently had a chance to use a beta release of Atlas on a real problem, and they are about to do a GA release, its ready for prime time.

Like most tools, it can produce masses of numbers and graphs, but compared to other storage analysis tools I've seen it goes further in three ways:

1) It collects critically important data that is not provided by the OS
2) It processes the data to tell you exactly what is wrong
3) It runs heuristics to tell you how to fix the problem

I wish more tools spent this much effort on solving the actual problem rather than making pretty graphs that only an expert would understand.

What we actually did was run the tool on a pre-production Oracle system using Veritas Filesystem and Volume Manager with Solaris on a SAN connected to a Hitachi storage array. Atlas starts off by looking at all the active processes on the system, and ignoring any that are not doing any I/O. It collects data on which files are being read or written by which process, and what the pattern and sizes are at the system call, file system and device level. You can also set the tool to focus on a set of devices, and gather information on the processes that actually talk to those devices.

Atlas immediately pointed out that two volumes had been concatenated to form a filesystem, and that 98% of the accesses were to one of the volumes. It recommended that the volumes be striped together for better overall performance.

It also pointed out that some of the I/O accesses were taking two seconds to complete at the filesystem level, but only two milliseconds at the device level. I guessed this was CPU starvation caused by fsflush running flat out on this machine which had over 50GB of RAM. Adding set autoup=600 to /etc/system and rebooting made the problem go away. We also saw this effect in the terminal window, where our typing would stop echoing for a few seconds every now and again. I've been told by Sun that the very latest patches finally fix fsflush so that it can't use a lot of CPU time, so large memory machines will finally work properly without needing this tweak.

Finally Atlas showed that the filesystem block size was set too small and Oracle was doing large reads that were being chopped into smaller reads by the filesystem layer before being sent to the device. It gave a specific recommendation for the block size that should be used. Reconfiguring the disks takes a long time to do, but we'll fix it before it goes into production.

We could have figured out the concatenation problem using iostat data, but the other two problems are normally invisible, and the topic of what filesystem block size to use can generate masses of discussion and confusion, so having "virtual Dave Fisk" tell you what blocksize to use can save a lot of time :-)