Sunday, July 29, 2012

Setting up vSphere Active / Active iSCSI connections to a NetApp FAS2040

I recently had the opportunity to architect a solution consisting of 3 vSphere 5 boxes connecting to a NetApp FAS2040.  Storage connectivity would be via iSCSI.  The storage network would be running off of 2 Cisco 2960G switches, soon to be replaced by stacked Cisco 3750’s. 

The requirements were stock standard, as high a throughput as possible, with as much redundancy as possible.  This meant going active active on the iSCSI links.  Here is how I did it.

NetApp FAS2040 Configuration

This little SAN has 8 1GB Ethernet ports.  Due to the fact that the Cisco 2960G switches does not support multi-link switch aggregation (this is where the 3750’s will come in) I had to come up with a simpler design – what NetApp terms a Single-Mode design.  My design allows for:

  • Two active connections to each controller, thus a total of four active sessions
  • Storage path HA
  • Load balancing across links
  • Uses vSphere storage MPIO as opposed to switch-side configuration

Virtual Interface (VIF) Configuration:

All Vif's are single-mode / active passive
Cont1_Vif01 - e0a/e0b (e0a will be active, connected to switch 1 / e0b passive connected to switch 2) IP – 192.168.1.1
Cont1_Vif02 - e0c/e0d (e0c will be active, connected to switch 2 / e0d passive connected to switch 1) IP – 192.168.2.1
Cont2_Vif01 - e0a/e0b (e0a will be passive, connected to switch 1 / e0b active connected to switch 2) IP – 192.168.1.2
Cont2_Vif02 - e0c/e0d (e0c will be passive, connected to switch 2 / e0d active connected to switch 1) IP – 192.168.2.2

This image, courtesy of NetApp, explains it infinitely better than my wall of text:-)

image

I also configured partner takeover for all VIF.  In case of controller failure it allows the remaining controller to take over the VIFs.

Ethernet Storage Network Configuration

On the storage network I had to configure 2 critical settings:

  • Spanning Tree Portfast
  • Jumbo Frames

When connecting ESX and NetApp storage arrays to Ethernet storage networks, NetApp highly recommends configuring the Ethernet ports to which these systems connect as RSTP edge ports.  This is done like so:

Switch2960(config)# interface gigabitethernet2/0/2
Switch2960(config-if)# spanning-tree portfast

Next up, Jumbo Frames:

Switch2960(config)# system mtu jumbo 9000
Switch2960(config)# exit
Switch2960# reload

vSphere Configuration

I am in love with vSphere 5, and one of the biggest reasons for that is the fact that a lot of the configuration parameters that used to be command-line only has been moved into the GUI.  Another reason is Multiple TCP Session Support for iSCSI.  This feature enables round robin load balancing using VMware native multipathing and requires a VMkernel port to be
defined for each physical adapter port assigned to iSCSI traffic.  That said, let’s get configuring:

  1. Open your vCenter Serve
  2. Select an ESXi host
  3. In the right pane, click the Configuration tab
  4. In the Hardware box, select Networking
  5. In the upper-right corner, click Add Networking to open the Add Network wizard
  6. Select the VMkernel radio button and click Next
  7. Configure the VMkernel by providing the required network information.  NetApp requires separate subnets for active/active iSCSI connections, therefore we will create two VMkernels, on the 192.168.1.x and 192.168.2.x subnets respectively.
  8. Configure each VMkernel to use a single active adapter that is not used by any other iSCSI VMkernel. Also, each VMkernel must not have any standby adapters. If using a single vSwitch, it is necessary to override the switch failover order for each VMkernel port used for iSCSI. There must be only one active vmnic, and all others should be assigned to unused
  9. The VMkernels created in the previous steps must be bound to the software iSCSI storage adapter. In the Hardware box for the selected ESXi server, select Storage Adapters.
  10. Right-click the iSCSI Software Adapter and select properties. The iSCSI Initiator Properties dialog box appears
  11. Click the Network Configuration tab
  12. In the top window, the VMkernel ports that are currently bound to the iSCSI software interface are listed
  13. To bind a new VMkernel port, click the Add button. A list of eligible VMkernel ports is displayed. If no eligible ports are displayed, make sure that the VMkernel ports have a 1:1 mapping to active vmnics as described earlier
  14. Select the desired VMkernel port and click OK.
  15. Click Close to close the dialog box
  16. At this point, the vSphere Client will recommend rescanning the iSCSI adapters. After doing this, go back into the Network Configuration tab to verify that the new VMkernel ports are shown as active, as per the image below.

image

Congratulations, you now have active / active, redundant iSCSI sessions into your NetApp SAN!

Saturday, July 14, 2012

Credibility, Ethics and Bias

 

I've been working in IT for the best part of a decade, but only got into blogging and the whole social media thing in the last year or so.  I really love doing what I do and sharing it with others, but in putting yourself out there you begin to realise how important ethics are.  There is absolutely no difference between me and the next blogger, apart from the quality of the content one puts up and credibility.

Then something dawned on me, credibility is not just something that should shine through in what you put out there for the public to consume, it is even more important to apply those principles in your day to day dealings.  It was about at that time when I realised that true credibility is something that is exceedingly rare in IT, in my experience.

In my universe, a very quick way to loose credibility is to shoot down and bad mouth a product, vendor or technology you know nothing about.  An example - I am in the somewhat unique situation where my job involves presales and architecting products from the two biggest storage vendors out there, namely EMC and NetApp.  As if that's not enough, I also do the HP EVA portfolio.  The storage field is hugely competitive, and this shows.  I take my job seriously, so I make it my business to know the products I work with as well as possible. 

For me it really is all about analysing the customers technical and business needs and consequently the application of the best technology for their given needs.  And believe me, there is enough key differentiators between the various vendors, that when combined with the customers budgetary requirements that you will be able to determine a best-fit solution, and not this one-size-fits-all that most vendors fixate on.  Unfortunately the amount of FUD and misinformation I've heard from people who really should know better is absolutely astounding.  It gets to the point where the vendors are *actively* just advancing their own best interests with the client and their interests a distant second (or maybe I'm just naive, and that is how its supposed to work?).

As if that's not bad enough, I also do the entire lifecycle of both vSphere and Hyper-V, from pre-sales through to implementing and supporting.  The amount of garbage I hear sprouted is enough to fill a landfill.  Admittedly most of it comes from the vSphere-supporting side of the fence, but the Microsoft partners are quickly catching up.  a Couple of examples I've heard is "Hyper-V does not do the equivalent of vMotion" or "the ESXi hypervisor is 50% faster then Hyper-V".  Complete and utter bollocks in other words.  As I said, the MS camp is quickly catching up and with the confidence and maturity that Hyper-V 3 will bring we'll see the MS guys giving as good as they get.

That being said, there will always be a bit of bias inherent in everyone.  You will develop bias through your career, naturally leaning towards the solutions that you sell and implement.  That is normal and there is nothing wrong with it.  By all means do challenge the opposition's claims, ask them to backup their statements, ask for facts, see through the normal sales BS and question their value propositions.

What is not right is the stuff I was talking about earlier.  At the risk of repeating myself we should all try and avoid spreading FUD intentionally.  Spreading it unintentionally is only slightly worse, because one should always verify claims before repeating it as gospel yourself.  If NetApp, for example, tells me they scored eleventy billion marks on some benchmark whilst EMC flunked out I will investigate.  EMC does knows a thing or two about storage - so there is bound to be a story behind the story.  Conversely, if I hear a EMC partner starting with "No one can touch our Avamar / DataDomain dedupe / our ease of management / etc" my BS detector goes into overdrive.

The ultimate loser here is the customer who gets bombarded with noise and misinformation from all sides, whose job hinges on making the correct decision, who ultimately needs to put his trust in a vendor who is more interested in pushing a brand or technology which might or might not solve a problem and who needs to explain when a solution does not deliver. 

We need to start putting the customer first in everything we do.  In the short term it might not seem the easy / profitable thing to do, but in the long term you will be rewarded.  Credibility is truly priceless, and once you give it up it is very, very difficult to regain. 

Do The Right Thing.

Sunday, June 3, 2012

Cluster Shared Volume stays in redirected mode

I recently had a perplexing problem on one of my lab servers, which took a lot of head-scratching to solve.  Fortunately I had some time to burn so I managed to get to the bottom of it.

Symptom

If I moved a disk or a CSV to a specific node in my Hyper-V failover cluster it would put the CSV in redirected mode and log the following to the System log

Log Name:      System
Source:        Microsoft-Windows-FailoverClustering
Event ID:      5125
Task Category: Cluster Shared Volume
Level:         Warning
Keywords:     
User:          SYSTEM
Description:
Cluster Shared Volume '\\?\Volume{0bf0b229-9b0e-11e1-8a3a-e4115ba98410}\' ('') has identified one or more active filter drivers on this device stack that could interfere with CSV operations. I/O access will be redirected to the storage device over the network through another Cluster node. This may result in degraded performance. Please contact the filter driver vendor to verify interoperability with Cluster Shared Volumes.
Active filter drivers found:
aksdf (Encryption)

Cause

After a fair bit of head-scratching, rolling back actions and research with Sysinternals Process Monitor I pinpointed the problem to NetApp Single Mailbox Restore for Exchange.  During installation it installs the aksdf.sys device driver.  A quick google search showed it to be a driver used for USB dongle licensing.  Weird, since SMBR does not require a dongle.  Anyhow, this device driver conflicts with the CSV and forces it to run in redirected mode

Solution

The solution is simple – navigate to the HKLM\SYSTEM\CurrentControlSet\Services\akdsf registry key and set the Start key to have a value of four (4), as per the below screenshot.

Image

This is not documented anywhere on the NetApp support site, so I will file a bug report.  In mitigation, I cannot see that one will actually run SMBR on one of your production cluster nodes.  Still, it should be trivial for NetApp to patch their installation routine to not install the aksdf.sys device driver.

Wednesday, May 23, 2012

NetApp Single Mailbox Restore (SMR) for Exchange for Virtualised Exchange Servers

NetApp Single Mailbox Restore for Exchange 2010 is, well, a snap to use when your Exchange server is running in the “NetApp way”.  What is the NetApp way you ask?  Well, in a nutshell, it is when you have your physical Exchange box hooked up to your SAN via iSCSI or FCP.  If you are virtualised then you’ll need to present your disks via RDM (vSphere) or pass-through if you live in MS land.

What I address here is the case where you have an Exchange server virtualised with Hyper-V, with your hard drives attached as VHD’s.  Even though this example uses Hyper-V, the principles are also applicable to a vSphere environment.

Mounting the NetApp Snapshot

  1. Open NetApp SnapDrive on a host connected to your Filer via either FCP or iSCSI
  2. Navigate to the Disks node and expand the LUN containing the VHD which in turn contains your Exchange DB’s.
    image
  3. Under Snapshot Copies, right-click point in time snapshot that you wish to restore and select Connect Disk
    image
  4. The Connect Disk Wizard will start.  Click Next
    image
  5. Select the appropriate snapshot and click next.
    image
  6. Click Next on the the “Important Properties…” screen (Don’t change anything here)
    image
  7. Set the LUN type as Dedicated and click Next
    image
  8. Assign a Drive Letter and click Next
    image
  9. Select your initiators and click Next
    image
  10. Select Manual on the Initiator Group Management Screen and click Next
    image
  11. Select the appropriate iGroup and click Next
    image
  12. Click Finish to complete the SnapDrive Connect Disk Wizard
    image

Your NetApp snapshot should now be mounted as a drive accessible through Windows Explorer, If you browse to it it should contain the VHD hosting your Exchange DB.  The next step is to mount the VHD so that it is accessible to SMR.

Mounting the VHD

  1. Open Server Manager.  Navigate to Storage – Disk Management.  Right-click Disk Management and click Attach VHD
    image
  2. Browse to the VHD from the previous section and click OK
    image

Your VHD will now be mounted with the next available drive letter and accessible via Windows Explorer.  The next and final step will be to mount our mailbox with SMR and get restoring!

Restoring with SMR

  1. Open Single Mailbox Recovery and click File – Open Source
    image
  2. Browse to your source EDB file, ignoring any warnings about missing log files (Hooray for application-aware snapshots!) and click OK
    image
  3. SMR will now process your database and allow you to restore a mailbox, folder or item to PST or an Exchange Server.
    image

Awesome, but once done we have to clean up after ourselves by dismounting the VHD and disconnecting the temporary NetApp SnapShot LUN.

Cleanup

  1. Open Server Manager. Navigate to Storage – Disk Management. Right-click Disk Management and click Detach VHD
    image
  2. Take care to *not* check the “Delete…” box and click OK
    image
  3. Open SnapDrive and go to the Disks node.  Right click your temporary SnapShot LUN and click Disconnect Disk.
    image

Thursday, May 10, 2012

Configuring NetApp SnapManager for Hyper-V (Part 2) – Adding your Hyper-V Failover Cluster


Part one of our little tutorial dealt with correctly setting up and sizing the Snapinfo LUN.  Part deux will show you how to add and configure your cluster for SnapManager for Hyper-V.  Let’s dive in.

Configuring a Hyper-V Failover Cluster

  1. Open up SnapManager for Hyper-V, click the Protection node – Hosts tab and click Add Host.  Enter your host name.  NB! Only enter the NetBIOS name, not the FQDN***
    image
  2. Click Next.  Answer Yes to the dialog box asking you to start the configuration wizard.image
  3. The configuration wizard will pop up
    image
  4. Click Next.  Enter the report path location (or choose the default).image
  5. Enter the correct notification settings for your environmentimage
  6. Click Next. Select your Snapinfo path.
    image
  7. Click Next. Admire the exquisitely formatted summary.image
  8. Click Finish.  The configuration wizard will now do the necessary to configure your Hyper-V failover cluster.
    image
  9. Once you click close you can start configuring your Hyper-V protection.

***If the Fully Qualified Domain Name (FQDN) is used, SMHV will not be able to recognize the name as a cluster. This is in view of the manner in which the Windows Failover Cluster (WFC) returns the cluster name through WMI calls. Consequently, the host will not be recognized by SMHV as a cluster and will fail to use a clustered LUN as the SnapInfo Directory Location.

Configuring NetApp SnapManager for Hyper-V (Part 1) – Creating the Snapinfo LUN

Simple as this sounds I found that the process is not as simple and as well documented as it could be, especially with regards to creating the clustered SnapInfo LUN and folders.  Consequently I decided to document it with (a first for this blog) screenshots.

I am going to assume that you have already hooked up your hosts to your NetApp system, and that you’ve installed SnapDrive and SnapManager for Hyper-V.

The steps, in a nutshell, are:

  1. Create the Snapinfo LUN
  2. Make the Snapinfo LUN a highly available clustered resource
  3. Configure SnapManager for Hyper-V

Creating the SnapInfo LUN

  1. Create a volume to host your Hyper-V SnapInfo LUN
  2. Open up Snapdrive on of your Hyper-V cluster nodes, go to the Disks node, and click Create Disk.  This launches the Create Disk Wizard.image
  3. Click Next.  Now highlight the volume you created in step 1, enter a LUN name and description:image
  4. Click Next.  Very Important – select Shared (Microsoft Cluster Services Only)image
  5. Click Next. The following list should list the active nodes in your Failover Cluster.image
  6. Click Next. Select the appropriate options and size for your environment***image
  7. Click Next. Select the initiators to be mapped to the LUNimage
  8. Click Next.  Select whether you want to manually select the igroups (collection of initiators) or whether you want the filer to do it automatically.image
  9. Click Next. Choose the option to create a new Cluster Group to host the LUNimage
  10. Click Next and click Finish to exit the wizard.

image

To recap, the above will:

  • Create a LUN on the volume of your choosing
  • Format the LUN with the NTFS filesystem
  • Add the disk to your Failover Cluster as part of a Cluster group
  • Assign a driveletter to the disk.

***SnapInfo LUN Size Provisioning:  The NetApp filer will store about 50KB metadata per VM per snapshot.  Due to the way Hyper-V snapshots work it will store two snaps per snapshot, therefore if we backup 20 VM’s once per day our sizing will be as follows:  20 * 50KB = 1MB * 2 = 2MB per day.  NetApp allows us to store 255 snapshots per volume so we should cater for 510 MB total.  I give it 10GB just because I can.  And because thin provisioning works.

Tuesday, May 8, 2012

Fixing NetApp SnapDrive error code 0x800706ba

I came across this when deploying SnapDrive on cluster nodes in a Windows 2008 R2 Failover Cluster.  Only SnapDrive running on the host owning the disk resource would list the connected LUNs.  On the other cluster nodes SnapDrive did not list any connected LUNs.  Event Viewer has the following to say:

Level: Warning
Source: Snapdrive
Event ID: 317
Description:  Failed to enumerate LUN.


Devicepath: '\\\mpio#disk&ven_netapp&prod_lun&rev_810a#1&7f6ac24&0&323766314c5d417548697a49#{53f56307-b6bf-11d0-94f2-00a0c91efb8b}'

Storage path: '/vol/vol_name_001/lun-name-001'

SCSI address: (3,0,0,7)

Error code: 0x800706ba

Error description: The RPC server is unavailable.

Background:

The issue is that SnapDrive on the non-owning hosts queries FCM to see which node is the owner.  It then tries to connect to Snapdrive on the owning node to retrieve the LUN and snapshot info (as opposed to directly from the filer).

Fix:

  1. Disable the windows Firewall (yeah right) –or
    Navigate to Control Panel > Windows Firewall > Allow a program through Windows
    Firewall > Exceptions.
  2. If you will be using HTTP or HTTPS, select the World Wide Web Services (HTTP) or Secure World Wide Web Services (HTTPS) checkboxes.
  3. Click Add program.
  4. Click Browse and browse to C:\Program Files\NetApp\SnapDrive\, or to wherever you
    installed SnapDrive if you did not use the default location.
  5. Select SWSvc.exe and click Open, then click OK in the Add a Program window and in the Windows Firewall Settings window.
  6. To verify that SWsvc.exe is in the list of inbound rules, in MMC, navigate to Windows Firewall > Inbound Rules.