Sunday, June 3, 2012

Cluster Shared Volume stays in redirected mode

I recently had a perplexing problem on one of my lab servers, which took a lot of head-scratching to solve.  Fortunately I had some time to burn so I managed to get to the bottom of it.

Symptom

If I moved a disk or a CSV to a specific node in my Hyper-V failover cluster it would put the CSV in redirected mode and log the following to the System log

Log Name:      System
Source:        Microsoft-Windows-FailoverClustering
Event ID:      5125
Task Category: Cluster Shared Volume
Level:         Warning
Keywords:     
User:          SYSTEM
Description:
Cluster Shared Volume '\\?\Volume{0bf0b229-9b0e-11e1-8a3a-e4115ba98410}\' ('') has identified one or more active filter drivers on this device stack that could interfere with CSV operations. I/O access will be redirected to the storage device over the network through another Cluster node. This may result in degraded performance. Please contact the filter driver vendor to verify interoperability with Cluster Shared Volumes.
Active filter drivers found:
aksdf (Encryption)

Cause

After a fair bit of head-scratching, rolling back actions and research with Sysinternals Process Monitor I pinpointed the problem to NetApp Single Mailbox Restore for Exchange.  During installation it installs the aksdf.sys device driver.  A quick google search showed it to be a driver used for USB dongle licensing.  Weird, since SMBR does not require a dongle.  Anyhow, this device driver conflicts with the CSV and forces it to run in redirected mode

Solution

The solution is simple – navigate to the HKLM\SYSTEM\CurrentControlSet\Services\akdsf registry key and set the Start key to have a value of four (4), as per the below screenshot.

Image

This is not documented anywhere on the NetApp support site, so I will file a bug report.  In mitigation, I cannot see that one will actually run SMBR on one of your production cluster nodes.  Still, it should be trivial for NetApp to patch their installation routine to not install the aksdf.sys device driver.

Wednesday, May 23, 2012

NetApp Single Mailbox Restore (SMR) for Exchange for Virtualised Exchange Servers

NetApp Single Mailbox Restore for Exchange 2010 is, well, a snap to use when your Exchange server is running in the “NetApp way”.  What is the NetApp way you ask?  Well, in a nutshell, it is when you have your physical Exchange box hooked up to your SAN via iSCSI or FCP.  If you are virtualised then you’ll need to present your disks via RDM (vSphere) or pass-through if you live in MS land.

What I address here is the case where you have an Exchange server virtualised with Hyper-V, with your hard drives attached as VHD’s.  Even though this example uses Hyper-V, the principles are also applicable to a vSphere environment.

Mounting the NetApp Snapshot

  1. Open NetApp SnapDrive on a host connected to your Filer via either FCP or iSCSI
  2. Navigate to the Disks node and expand the LUN containing the VHD which in turn contains your Exchange DB’s.
    image
  3. Under Snapshot Copies, right-click point in time snapshot that you wish to restore and select Connect Disk
    image
  4. The Connect Disk Wizard will start.  Click Next
    image
  5. Select the appropriate snapshot and click next.
    image
  6. Click Next on the the “Important Properties…” screen (Don’t change anything here)
    image
  7. Set the LUN type as Dedicated and click Next
    image
  8. Assign a Drive Letter and click Next
    image
  9. Select your initiators and click Next
    image
  10. Select Manual on the Initiator Group Management Screen and click Next
    image
  11. Select the appropriate iGroup and click Next
    image
  12. Click Finish to complete the SnapDrive Connect Disk Wizard
    image

Your NetApp snapshot should now be mounted as a drive accessible through Windows Explorer, If you browse to it it should contain the VHD hosting your Exchange DB.  The next step is to mount the VHD so that it is accessible to SMR.

Mounting the VHD

  1. Open Server Manager.  Navigate to Storage – Disk Management.  Right-click Disk Management and click Attach VHD
    image
  2. Browse to the VHD from the previous section and click OK
    image

Your VHD will now be mounted with the next available drive letter and accessible via Windows Explorer.  The next and final step will be to mount our mailbox with SMR and get restoring!

Restoring with SMR

  1. Open Single Mailbox Recovery and click File – Open Source
    image
  2. Browse to your source EDB file, ignoring any warnings about missing log files (Hooray for application-aware snapshots!) and click OK
    image
  3. SMR will now process your database and allow you to restore a mailbox, folder or item to PST or an Exchange Server.
    image

Awesome, but once done we have to clean up after ourselves by dismounting the VHD and disconnecting the temporary NetApp SnapShot LUN.

Cleanup

  1. Open Server Manager. Navigate to Storage – Disk Management. Right-click Disk Management and click Detach VHD
    image
  2. Take care to *not* check the “Delete…” box and click OK
    image
  3. Open SnapDrive and go to the Disks node.  Right click your temporary SnapShot LUN and click Disconnect Disk.
    image

Thursday, May 10, 2012

Configuring NetApp SnapManager for Hyper-V (Part 2) – Adding your Hyper-V Failover Cluster


Part one of our little tutorial dealt with correctly setting up and sizing the Snapinfo LUN.  Part deux will show you how to add and configure your cluster for SnapManager for Hyper-V.  Let’s dive in.

Configuring a Hyper-V Failover Cluster

  1. Open up SnapManager for Hyper-V, click the Protection node – Hosts tab and click Add Host.  Enter your host name.  NB! Only enter the NetBIOS name, not the FQDN***
    image
  2. Click Next.  Answer Yes to the dialog box asking you to start the configuration wizard.image
  3. The configuration wizard will pop up
    image
  4. Click Next.  Enter the report path location (or choose the default).image
  5. Enter the correct notification settings for your environmentimage
  6. Click Next. Select your Snapinfo path.
    image
  7. Click Next. Admire the exquisitely formatted summary.image
  8. Click Finish.  The configuration wizard will now do the necessary to configure your Hyper-V failover cluster.
    image
  9. Once you click close you can start configuring your Hyper-V protection.

***If the Fully Qualified Domain Name (FQDN) is used, SMHV will not be able to recognize the name as a cluster. This is in view of the manner in which the Windows Failover Cluster (WFC) returns the cluster name through WMI calls. Consequently, the host will not be recognized by SMHV as a cluster and will fail to use a clustered LUN as the SnapInfo Directory Location.

Configuring NetApp SnapManager for Hyper-V (Part 1) – Creating the Snapinfo LUN

Simple as this sounds I found that the process is not as simple and as well documented as it could be, especially with regards to creating the clustered SnapInfo LUN and folders.  Consequently I decided to document it with (a first for this blog) screenshots.

I am going to assume that you have already hooked up your hosts to your NetApp system, and that you’ve installed SnapDrive and SnapManager for Hyper-V.

The steps, in a nutshell, are:

  1. Create the Snapinfo LUN
  2. Make the Snapinfo LUN a highly available clustered resource
  3. Configure SnapManager for Hyper-V

Creating the SnapInfo LUN

  1. Create a volume to host your Hyper-V SnapInfo LUN
  2. Open up Snapdrive on of your Hyper-V cluster nodes, go to the Disks node, and click Create Disk.  This launches the Create Disk Wizard.image
  3. Click Next.  Now highlight the volume you created in step 1, enter a LUN name and description:image
  4. Click Next.  Very Important – select Shared (Microsoft Cluster Services Only)image
  5. Click Next. The following list should list the active nodes in your Failover Cluster.image
  6. Click Next. Select the appropriate options and size for your environment***image
  7. Click Next. Select the initiators to be mapped to the LUNimage
  8. Click Next.  Select whether you want to manually select the igroups (collection of initiators) or whether you want the filer to do it automatically.image
  9. Click Next. Choose the option to create a new Cluster Group to host the LUNimage
  10. Click Next and click Finish to exit the wizard.

image

To recap, the above will:

  • Create a LUN on the volume of your choosing
  • Format the LUN with the NTFS filesystem
  • Add the disk to your Failover Cluster as part of a Cluster group
  • Assign a driveletter to the disk.

***SnapInfo LUN Size Provisioning:  The NetApp filer will store about 50KB metadata per VM per snapshot.  Due to the way Hyper-V snapshots work it will store two snaps per snapshot, therefore if we backup 20 VM’s once per day our sizing will be as follows:  20 * 50KB = 1MB * 2 = 2MB per day.  NetApp allows us to store 255 snapshots per volume so we should cater for 510 MB total.  I give it 10GB just because I can.  And because thin provisioning works.

Tuesday, May 8, 2012

Fixing NetApp SnapDrive error code 0x800706ba

I came across this when deploying SnapDrive on cluster nodes in a Windows 2008 R2 Failover Cluster.  Only SnapDrive running on the host owning the disk resource would list the connected LUNs.  On the other cluster nodes SnapDrive did not list any connected LUNs.  Event Viewer has the following to say:

Level: Warning
Source: Snapdrive
Event ID: 317
Description:  Failed to enumerate LUN.


Devicepath: '\\\mpio#disk&ven_netapp&prod_lun&rev_810a#1&7f6ac24&0&323766314c5d417548697a49#{53f56307-b6bf-11d0-94f2-00a0c91efb8b}'

Storage path: '/vol/vol_name_001/lun-name-001'

SCSI address: (3,0,0,7)

Error code: 0x800706ba

Error description: The RPC server is unavailable.

Background:

The issue is that SnapDrive on the non-owning hosts queries FCM to see which node is the owner.  It then tries to connect to Snapdrive on the owning node to retrieve the LUN and snapshot info (as opposed to directly from the filer).

Fix:

  1. Disable the windows Firewall (yeah right) –or
    Navigate to Control Panel > Windows Firewall > Allow a program through Windows
    Firewall > Exceptions.
  2. If you will be using HTTP or HTTPS, select the World Wide Web Services (HTTP) or Secure World Wide Web Services (HTTPS) checkboxes.
  3. Click Add program.
  4. Click Browse and browse to C:\Program Files\NetApp\SnapDrive\, or to wherever you
    installed SnapDrive if you did not use the default location.
  5. Select SWSvc.exe and click Open, then click OK in the Add a Program window and in the Windows Firewall Settings window.
  6. To verify that SWsvc.exe is in the list of inbound rules, in MMC, navigate to Windows Firewall > Inbound Rules.

Friday, April 27, 2012

SQL Development and Testing magic, brought to you by NetApp

I started supporting and deploying NetApp storage systems about 5 months ago.  I still am *very* far from being an authority on the subject, but I have been playing around with the kit and am learning an immense amount daily.  That said, I never seize to be amazed at how trivially easy NetApp technology make traditionally time-consuming and difficult tasks.  Take the following situation:  You have a big old production SQL DB, and your developers need to test against an up-to-date instance of it.  I bet the tired old way that your business is doing it is by taking a 500GB backup, pushing it across the network, or sneaker-netting it via USB or even worse restoring from tape (blegh!).

No more.  With a NetApp SAN I can do this in exactly 23 seconds (and this on a low-end FAS2040!).  It costs me zero storage*** on my SAN and all it takes is you putting on your PowerShell hat and typing one single line.  The only assumption I make here is that you are running SnapManager for Microsoft SQL Server on your SQL box (and you should!).  This gives us access to all the PowerShell goodness that makes the magic happen.  The following command (watch for wrappage):  

clone-backup -svr prod-sql-001 -Database big_prod_db -TargetDatabase yourdatabase -TargetServerInstance dev-sql-001 Backup sqlsnap__big_prod_db__snapshots

  1. What is happening behind the scenes, you ask?  Quite simple actually:
  2. It created a FlexClone of the Snapshot of the production store
  3. Created a DB on dev-sql-001
  4. Mounted the FlexClone as a mountpoint
  5. Attached the files to the database
  6. Made toast and coffee

But that's not all.  This example brought our development box up to the same point in time as production, but you can also specify a previous a point in time to go back to.  Of course you need to have snapshots of your chosen point in time.

Wow.

***It will start consuming space once you start making changes in dev, because you are diverging from the cloned snapshot.

Wednesday, April 11, 2012

Issues when upgrading to Veeam Backup & Replication v6

Let me start off by saying that Veeam Backup & Replication (VBR) is one of the most awesome pieces of software I know of!  As with any piece of software there might be the occasional bug or issue that needs to be worked around.  I’ve upgraded a couple of clients to VBR already, and I’ve consistently run into two little niggles.  These are:

Backup of vCenter SQL DB fails when using VBR v6

This is actually detailed on the Veeam forums here.

What it boils down to is that when using VBR for an application aware backup of the VM which hosts the vCenter SQL DB, the backup will fail with a VSSControl: Failed to freeze guest, wait timeout error.  This occurs because VBR has to communicate with vCenter to create a snapshot, but at the same time the vCenter SQL DB is frozen because of the VSS snapshot

To solve this problem, add the IP-address of the ESX host which runs the vCenter Server SQL database to the Veeam console (the servers list in the left panel). Then adjust the job and select the virtual machine with the SQL database from the ESX host instead of via the vCenter Server. Veeam B&R will then use the host for communication with the VM to do the VSS snapshot.

This will, of course, cause your replication job to fail should the vCenter VM move to another host.  We can work around this by using DRS Host Affinity rules to tie your vCenter VM to a physical box.

Replication job fails – cannot connect to port 2500

When VBR does a replication job, it might fail with “Creating snapshot Cannot connect to server [x.x.x.x:2500]”

In my case it was always happening when the destination server was running ESX (not ESXi) v4.x.  This is easily resolved by logging onto the target server via SSH and issuing the following commands:

  1. esxcfg-firewall -openport 2500:2510,tcp,in,VeeamSCP
  2. service mgmt-vmware restart

Easy enough!