Friday, 27 April 2012


What is Data ?

Data is a collection of raw facts from which conclusions may be drawn.Handwritten letters, a printed book, a family photograph, a movie on video tape, printed and duly signed copies of mortgage papers, a bank’s ledgers, and an account holder’s passbooks are all examples of data.Today, the same data can be converted into more convenient forms such as an
e‑mail message, an e-book, a bitmapped image, or a digital movie. This data can be generated using a computer and stored in strings of 0s and 1s, as shown in Figure




Types Of Data:

Data can be classified as structured or unstructured based on how it is stored and managed. Structured data is organized in rows and columns in a rigidly defined format so that applications can retrieve and process it efficiently. Structured data is typically stored using a database management system (DBMS).Data is unstructured if its elements cannot be stored in rows and columns,and is therefore difficult to query and retrieve by business applications. For example, customer contacts may be stored in various forms such as sticky notes, e-mail messages, business cards, or even digital format files such as .doc, .txt,
and .pdf. Due its unstructured nature, it is difficult to retrieve using a customer relationship management   application. Unstructured data may not have the required components to identify itself uniquely for any type of processing or     interpretation. Businesses are primarily concerned with managing unstructured data because over 80 percent of enterprise data is unstructured and requires significant storage space and effort to manage.





What is Storage?

Data created by individuals or businesses must be stored so that it is easily accessible for further processing. In a computing environment, devices designed for storing data are termed storage devices or simply storage. The type of storage used varies based on the type of data and the rate at which it is created and used. Devices such as  memory in a cell phone or digital camera, DVDs, CD-ROMs, and hard disks in personal computers are examples of storage devices. Businesses have several options available for storing data including internal hard disks, external disk arrays and tapes.

Evolution of Storage Technology and Architecture

Organizations had centralized computers (mainframe) and information storage devices (tape reels and disk packs) in their data center. The evolution of open systems and the affordability and ease of deployment that they offer made it possible for business units/departments to have their own servers and storage. In earlier implementations of open systems, the storage was typically internal to the server.
Redundant Array of Independent Disks (RAID): This technology was developed to address the cost, performance, and availability requirements of data. It continues to evolve today and is used in all storage architectures such as DAS, SAN, and so on.
Direct-attached storage (DAS): This type of storage connects directly to a server (host) or a group of servers in a cluster. Storage can be either internal or external to the server. External DAS alleviated the challenges of limited internal storage capacity.
Storage area network (SAN): This is a dedicated, high-performance Fibre Channel (FC) network to facilitate block-level communication between servers and storage. Storage is partitioned and assigned to a server for accessing its data. SAN offers scalability, availability, performance, and cost benefits compared to DAS.

Network-attached storage (NAS): This is dedicated storage for file serving applications. Unlike a SAN, it connects to an existing communication network (LAN) and provides file access to heterogeneous clients. Because it is purposely built for providing storage to file server applications, it offers higher scalability, availability, performance, and cost benefits compared
to general purpose file servers.

Internet Protocol SAN (IP-SAN): One of the latest evolutions in storage architecture, IP-SAN is a convergence of technologies used in SAN and NAS. IP-SAN provides block-level communication across a local or wide area network (LAN or WAN), resulting in greater consolidation and availability of data.



What is Data Center ?

Organizations maintain data centers to provide centralized data processing capabilities across the enterprise. Data centers store and manage large amounts of mission-critical data. The data center infrastructure includes computers, storage systems, network devices, dedicated power backups, and environmental controls (such as air conditioning and fire suppression).
Large organizations often maintain more than one data center to distribute data processing workloads and provide backups in the event of a disaster. The storage requirements of a data center are met by a combination of various storage architectures.

Application: An application is a computer program that provides the logic for computing operations. Applications, such as an order processing system, can be layered on a database, which in turn uses operating system services to perform read/write operations to storage devices.

Database: More commonly, a database management system (DBMS) provides a structured way to store data in logically organized tables that are interrelated. A DBMS optimizes the storage and retrieval of data.

Server and operating system: A computing platform that runs applications and databases.

Network: A data path that facilitates communication between clients and servers or between servers and storage.

Storage array: A device that stores data persistently for subsequent use.  These core elements are typically viewed and managed as separate entities, but all the elements must work together to address data processing requirements.



Dear Readers,


I am going to write on Storage networks and This course is targeted towards developers, integrators, managers and others with a need for a comprehensive, in-depth understanding of the Storage technology.An understanding of current computer interfaces or networks is desirable, although not absolutely necessary.






Storage networks course:


  1. Introduction to Storage

  1. Storage Topologies

  1. Interface Protocols

  1. RAID

  1. Network Attached Storage

  1. Fibre Channel protocol

  1. Storage Area Networks

  1. IP SAN –ISCSI

  1. Business continuity with Backup and replication

  1. Introduction to File system

  1. Era of Virtualization

  1. Data Deduplication
      

Tuesday, 13 March 2012

Big Data Market to Grow to $16.9 Billion by 2015: IDC


A new IDC study says the market for big data technology and services will grow from $3.2 billion in 2010 to $16.9 billion in 2015.

Market research firm IDC has released a new forecast that shows the big data market is expected to grow from $3.2 billion in 2010 to $16.9 billion in 2015.
Big data is a new frontier in IT where data sets can grow so large that they become awkward to work with using traditional database management tools. Thus, the need for new and more tools, frameworks, hardware, software and services to handle this emerging issue represents a huge market opportunity.
IDC projects that the increased investment needed to support big data represents a compound annual growth rate (CAGR) of 40 percent, or about seven times that of the overall information and communications technology (ICT) market.
According to IBM, everyday business and consumer life creates 2.5 quintillion bytes of data per day—so much that 90 percent of the data in the world today has been created in the last two years alone. This data comes from everywhere: from sensors used to gather climate information, posts to social media sites, digital pictures and videos posted online, transaction records of online purchases, and from cell phone GPS signals, to name a few.
"The big data market is expanding rapidly as large IT companies and startups vie for customers and market share," Dan Vesset, program vice president for Business Analytics Solutions at IDC, said in a statement. "For technology buyers, opportunities exist to use big data technology to improve operational efficiency and to drive innovation. Use cases are already present across industries and geographic regions.
Moreover, while the five-year CAGR for the worldwide market is expected to be nearly 40 percent, the growth of individual segments varies from 27.3 percent for servers and 34.2 percent for software to 61.4 percent for storage.
"There are also big data opportunities for both large IT vendors and startups," Vesset said. "Major IT vendors are offering both database solutions and configurations supporting big data by evolving their own products as well as by acquisition. At the same time, more than half a billion dollars in venture capital has been invested in new big data technology."
Additionally, IDC said the growth in appliances, cloud computing and outsourcing deals for big data technology will likely mean that over time end users will pay increasingly less attention to technology capabilities and will focus instead on the business value arguments. System performance, availability, security and manageability will all matter greatly. However, how they are achieved will be less of a point for differentiation among vendors.
IDC also said today there is a shortage of trained big data technology experts, in addition to a shortage of analytics experts. This labor supply constraint will act as an inhibitor of adoption and use of big data technologies, and it will also encourage vendors to deliver big data technologies as cloud-based solutions.
"While software and services make up the bulk of the market opportunity through 2015, infrastructure technology for big data deployments is expected to grow slightly faster at a 44 percent CAGR, Benjamin Woo, program vice president for Storage Systems at IDC, said in a statement. “Storage, in particular, shows the strongest growth opportunity, growing at 61.4 percent CAGR through 2015. The significant growth rate in revenue is underscored by the large number of new open-source projects that drive infrastructure investments."
IDC's methodology for sizing the big data technology and services market includes the evaluation of current and expected deployments that follow one of the following three scenarios:
  • Deployments where the data collected is over 100 terabytes (TB). IDC is using data collected, not stored, to account for the use of in-memory technology where data may not be stored on a disk.
  • Deployments of ultra-high-speed messaging technology for real-time, streaming data capture and monitoring. This scenario represents big data in motion as opposed to big data at rest.
  • Deployments where the data sets may not be very large today, but are growing very rapidly at a rate of 60 percent or more annually.
Additionally, IDC requires that in each of these three scenarios, the technology is deployed on scale-out infrastructure and deployments that include either two or more data types or data sources or those that include high-speed data sources, such as click-stream tracking or monitoring of machine-generated data.

Cisco UCS M3 and Intel Romley


Cisco UCS M3 and Intel Romley


Today Cisco officially unveils the next generation of UCS servers.  These go along with the new announcement by Intel this week of their “Romley” platform that use the E5-2600 family of CPUs.  Let’s take a look at what’s new, changed, and where that puts us.

First, Let’s Look at Intel’s Announcement

These new blades are based on the new E5-2600 series of CPUs.  The codename for this whole platform is/was Romley.  The E5-2600 CPUs were code named Sandy Bridge.  If you’ve seen those names floating around now you know how it all comes together.  These CPUs replace the Xeon 5600 line that has been the basis for the B200 and C200 series of servers from Cisco.  Where as the 5600 line topped out at 6 physical cores per CPU the new E5-2600 line goes to 8 cores.  Just like the 5600 line they have HyperThreading so they can execute up to 2 threads per clock cycle.  This means we’ll now have dual-socket servers with 16 physical cores capable of executing 32 threads.
Memory is usually the constraint in a virtual environment, and that’s where we play most often.  These new E5s can do 24 DIMM slots in a single dual-socket server.  Combine those slots with the newly available 32GB DIMMs and you have a real small powerhouse that can do 16 cores and 768GB of RAM.  Along with this you’re going to see much higher memory throughput, as well.  Maximum memory speed goes up to 1600MHz with up to two DIMMs per bank.  When you go to 3 the speed drops to 1066MHz.  The QPI (QuickPath Interconnect), the direct path between CPUs, is now faster and there are now two of them.  Previously it was 6.4GT/s and now it is 8GT/s.
Along with fast RAM you also get more maximum cache.  The higher end chips go to 20MB of onboard cache.  You can see how these are really starting to blur lines with the E7 line of processors.
Another enhancement is what is referred to as Turbo 2.0.  The previous generation of chips would raise the clock speed of the CPU up when it could.  That was when the thermal load allowed for faster speed and usually when not all cores were utilized.  If you had a 6 core CPU but only 2 cores were active and the temperature of the chip had some headroom the system would “turbo” up the CPU in 133MHz increments.  These new CPUs are more aggressive and will actually go above the TDP (Thermal Design Power) for a brief period of time.  Basically, you should see turbo mode more often which leads to faster processing.
I was going to put a table here of the new CPUs but I’ll let the fine people at Anandtech do that for me here.  They also have a more in-depth overview of the new offerings.
The bottom line is that these CPUs are fast…of course they are faster than the 5600s they replace but it’s not just more cores and faster clockspeed.  In fact, clock speed is stagnant or getting a little slower yet these CPUs still get more done in a cycle than their predecessors.  When looking and comparing you need to go find good benchmarks or CPU leveling scales to make accurate estimations for sizing.  Some benchmarks I’ve seen show the same clockspeed CPU on the 5600 and E5 lines and the E5 almost 50% faster.  That’s impressive.
Oh yeah..one more thing…  Eventually Intel will have a line of E5 CPUs that are quad-socket capable.  Consider that a less expensive quad-socket option than the E7s.

What is Cisco Doing With These?

First, Cisco is not releasing M3 versions of all servers.  The new E5-2600s do not replace the E7 line of CPUs so the B230M2 will continue to be offered in its current form for a while.  They are releasing a B200M3, C220M3, and C240M3.
Now let’s start with the B-series enhancements.
So along with the new CPU and chipset, what else is new for the M3 series of blades?  A number of things, especially around I/O.  Recently we saw the introduction of new second generation interconnects (6248UP and soon the 6296UP) along with the new FEX (Fabric Extenders), the 2208XPs with 8 ports each.  To help drive that sort of I/O from the blades Cisco released…well..announced…the VIC1280 mezzanine card.  This card gives a blade eight 10Gb ports for a total I/o throughput of 80Gb.  Along with the new M3 blades is the introduction of the VIC1240.
This new module provides 40Gb of throughput using two 10Gb ports to each fabric and lets you create up to 256 virtual devices (vNICs and vHBAs).  This will basically replace the M81KR (Palo) adapter that has been our good friend for a while.  But wait!  There’s more.  Along with the VIC1240 you can add an optional module, the Port Expander Card.  This takes the VIC1240 up to 80Gb of throughput, much like a VIC1280…but you can add this optional expansion at any time.
These modules will be known as MLoM, Modular LAN on Motherboard.  Why not just LoM?  Because they are optional.  There was talk earlier about building the VIC1240 on the board but that’s not the case in the final version.  These new MLoM adapters fit on a connector on the blade.
The older generation mezzanine adapters will NOT fit these new blades.  That also means the other adapters, like Qlogic and Emulex CNAs, will be revised to a new version as well.  Yes, they are still available just like they were before.  But….I see two slots there.  What’s the other one for?  3rd party modules.  What sort of modules?  Well…I get asked all the time for other things on UCS blades like GPUs for rendering…Fusion I/O type caching cards…VFCache from EMC was recently announced and needs a PCIe type card.  I suspect will see announcements about these other modules soon.  And yes, you absolutely can do the VIC card and a 3rd party card at the same time.
Note:  If you add the Port Expansion to the VIC1240 that means you can’t use a 3rd party module as the Port Expansion uses that mezzanine card slot.
What else about the new B200M3?
We already talked about the mezzanine adapters.  There is an internal USB 2.0 port.  You can see the 24 DIMM slots in the picture.  It still retains the two hot swap SAS drives in the front but now you have the option for SSDs on the B200.  The M3 blades now also have two internal SD card slots.  Eventually you’ll be able to boot from these but that won’t be available right away.  My guess is that we’re waiting on an updated UCS Manager since there is no boot option for that currently.  This new feature is called “Cisco Flexible Flash”.  Pretty exciting name for two SD slots.

A Word About I/O

Along with these new blades another FEX for the UCS chassis was released, the 2204XP.  It is the second generation 4-port FEX.  To take advantage of all features of the VIC1240 and the Port Expander you’ll need to upgrade from the 2104XP to the 2204XP or 2208XP.   I suspect we’ll see some good upgrade offers for these, but I don’t know for sure.  That does explain why the 2208XP has been the FEX included in the UCS bundles lately.  Cisco should be releasing a support and features matrix showing the VICs and the FEX options and how they work together.

 And  Now the C-Series Offerings

I won’t lie…the vast majority of UCS business that I’m involved in is B-series but Cisco’s C-series rackmount servers continue to evolve.  Along with the new B200M3 Cisco is releasing the C220M3 and C240M3.
The picture above is the new C220M3 and I have to say, the servers are at least starting to look nicer!  This is your standard 1U server that can do two CPUs for a total of 16 physical cores and 768GB of RAM in 24 DIMM slots.  It has two PCIe slots.
That is the C240M3.  Can you guess what its claim to fame is?  If you said disk density you were right!  Since it’s based on the E5-2600 line it is also dual socket and can go up to 768GB of RAM.  It has 5 PCIe slots and can hold 24 SAS drives in 2U.
Cisco has recently changed the architecture for how you can connect and manage the C-series servers under UCS Manager.  I’ll do another post on that soon, but the good news here is that now all C-series will support being used in that type of configuration.  It’s not something that has been popular in the past due to the complexity but I’m interested to see how the new offering does.  It use a lot less ports and makes much more sense.

Dual BIOS Support

This falls under both B and C-series and is worth mentioning.  All the new M3 systems have dual BIOS capability so that they can recover from a corrupt BIOS image.  Previously this would require RMAing the system, but that’s no longer the case.

Jason’s Opinion

You can probably stop reading here…  My thoughts are simple.  The only real surprise here is the updated I/O capability and it’s not totally unexpected.  Cisco tipped their hand a bit with the 2208XP FEX a while back.  These are all very good evolutions, especially the option for 3rd party mezzanine cards.  That’s something that customers are starting to ask about and with EMC’s announcement of VFCache and their involvement with VCE there had to be a solution for that.  You’ll find I’m a big fan of options and these new offerings give you options.  Intel has done a great job building a platform that Cisco then takes to the next level.
As Intel expands out the E5 line you’ll see Cisco revise the blade offerings but for now we get a good taste of what they are thinking.  I am looking forward to the quad-socket E5 CPUs as well as the price of RAM continuing to drop so that 16GB DIMMs get even cheaper and one day those new 32GB DIMMs become affordable.  768GB in the standard workhorse B200M3 blade!  CRAZY!
One advantage that Cisco has had for a while now is their extended memory support using custom developed silicone.  They could do more memory density than others while also keeping the memory speed high at 1333MHz or 1066MHz.  It’s been a serious differentiator for them.  With the new Romley platform the playing field is now level.  If you look at the B200M3 you’ll see there just isn’t room for more DIMM slots and you can’t go bigger on DIMMs than 32GB (at least for now).
But what about the B230 and the E7 line of CPUs?  Aren’t these encroaching on that territory?  Yes and no.  The B230 still offers 10-core E7 CPUs to give you 20 physical cores on a single half-width blade.  Memory support on those blades is still very good and enough for almost everyone.  The E7 also has other features that the E5 line does not.  One example is RAS (Reliability, Availability, Serviceability).  I need to do a good writeup on RAS as many people have no idea it’s in there but it’s a key feature on those CPUs.  For example, if the CPU detects an unrecoverable memory error it can tell vSphere to kill the VM using that RAM but leave the others untouched.  In a system without RAS that would cause a vSphere host to purple screen and take all VMs down.  Very important for business critical applications.  That’s why the B230 and B440 are still very relevant today and I’d have no problem deploying more.
These new offerings can be ordered but probably won’t ship for at least a month.  Work with us (preferably!) or your partner for pricing.

Tuesday, 28 February 2012

Demartek Storage Networking Interface Comparison


Demartek Storage Networking Interface Comparison

Updated 22 February 2012
By , Demartek President
Because of the number of storage interface types and related technologies that are used for storage devices, we have compiled this summary document providing some basic information for each of the interfaces. This document will be updated periodically. This document may become larger over time.Contact us if you’d like to see additional information in this document.
The interface types listed here are known as “block” interfaces, meaning that they provide an interface for “block” reads and writes. They simply provide a conduit for blocks of data to be read and written, without regard to file systems, file names or any other knowledge of the data in the blocks. The host requesting the block access provides a starting address and number of blocks to read or write.
We are producing deployment guides for some of the interfaces described in this document. The Demartek iSCSI Deployment Guide 2011 is now available. More Read here